Alex Turner, also known online as “TurnTrout,” is a researcher specializing in artificial intelligence safety and AI alignment. His research spans reinforcement learning, theories of agent behavior, machine learning interpretability, and the internal representations of large language models. Unlike researchers primarily concerned with improving model capabilities, Turner focuses on a more fundamental question with far-reaching implications: as AI systems become increasingly powerful and capable of autonomously completing complex tasks, how can humans understand, predict, and control their behavior so that these systems continue to act in accordance with what people genuinely want?
On his personal website, Turner presents his research as an effort to understand and control advanced AI systems. The “alignment” problem he studies is not merely about preventing chatbots from producing harmful content, nor is it limited to imposing superficial behavioral rules on models. At a deeper level, it concerns questions such as: Why does a system develop a particular objective? How will it interpret the rewards it received during training when placed in a new environment? Once a system acquires the ability to plan, might it seek resources, expand its influence, or even resist external intervention in order to accomplish its objectives? These questions form a central thread throughout Turner’s research career.
Formal Research on Power-Seeking
One of Turner’s most representative early research areas was the tendency of AI agents to engage in “power-seeking.” A recurring concern in AI safety is that even when a system’s final objective appears harmless, it may still treat acquiring resources, preserving its own operation, controlling its environment, or preventing shutdown as useful means of achieving that objective. For example, a system tasked only with increasing a production metric might discover that gaining access to more computing resources, permissions, and decision-making authority improves its chances of success.
Historically, this possibility was often discussed through thought experiments or intuitive arguments. Turner sought to explain it mathematically by identifying the conditions under which optimal policies tend to preserve more options or move agents into positions from which they can influence a greater number of future states. Research to which he contributed, including work on the proposition that “optimal policies tend to seek power,” connects power with an agent’s ability to control the future, preserve access to possible states, and accomplish a wide range of potential objectives. The importance of this work does not lie in claiming that every AI system will inevitably seek power. Instead, it transforms a vague safety concern into a research problem that can be formally defined, analyzed, and tested.
This formal work also helps clarify a common misunderstanding in public discussions of AI. Dangerous AI behavior does not necessarily arise from human-like ambition, desire, or malice. A system does not need to “want to rule” in order to behave in ways that increase its control. Such behavior may instead emerge from the structure of its objective and the optimization pressures created by its environment. In other words, the risk may originate from structural relationships between goals and environments rather than from any personified, malicious motive within the machine.
Turner has also studied methods for limiting the excessive side effects that agents may impose on their environments. One representative approach, known as “Attainable Utility Preservation,” attempts to penalize actions that substantially reduce the range of future possibilities. For example, a cleaning robot should not destroy objects in a room simply to complete its task more quickly, because such destruction would permanently eliminate many other objectives that might otherwise remain achievable. Although methods of this kind cannot solve the entire alignment problem on their own, they provide theoretical tools for measuring side effects, preserving reversibility, and designing more cautious agents.
From External Behavior to Internal Mechanisms
As large language models developed rapidly, Turner’s research expanded from abstract theories of agent behavior to the internal mechanisms of neural networks. He began investigating questions such as: How do models internally represent concepts, intentions, and behavioral tendencies? Can researchers directly identify these representations and use them to alter model behavior?
Conventional methods for controlling models generally rely on prompt engineering, supervised fine-tuning, or reinforcement learning from human feedback. These approaches influence models mainly through their inputs and outputs, and they often require substantial amounts of data and computation. “Activation engineering,” an area Turner has helped advance, offers another route: directly analyzing and modifying the internal activations produced while a neural network is operating.
In related research, investigators can compare the internal states generated when a model processes two sets of texts with opposing meanings, such as honesty and deception, positivity and negativity, or acceptance and refusal. From these comparisons, they can extract a direction corresponding to a particular semantic or behavioral feature. Adding or subtracting this direction while the model generates text may then consistently strengthen or weaken the associated behavior. Such techniques are often described as “activation addition” or “representation engineering.” They suggest that at least some high-level concepts in large language models are not entirely distributed and incomprehensible. Instead, they may be encoded through relatively regular geometric structures within activation space.
Turner’s work on refusal mechanisms further demonstrates the significance of this approach. Research has found that, in some cases, a language model’s refusal to comply with harmful requests is closely associated with a particular direction in its internal representation space. Strengthening this direction can increase the model’s tendency to refuse, while weakening or removing it may allow the model to bypass its existing safety training.
This finding has two important implications. On the one hand, it shows that some aspects of a model’s safety behavior can be located, interpreted, and manipulated. On the other hand, it exposes the fragility of existing safeguards. If a complex refusal mechanism depends too heavily on a small number of internal features, a model that appears safe on the surface may not possess a stable or comprehensive understanding of safety.
Shard Theory and the Formation of Values
Turner has also helped develop and discuss “Shard Theory,” an approach intended to explain how agents trained through reinforcement learning gradually acquire internal structures resembling values, preferences, or motivations. A “shard” is not a complete objective programmed into a system in advance. Instead, it is a localized decision-making tendency formed through particular training experiences. As the system accumulates experience, these tendencies may compete with one another, combine, and gradually generalize, eventually influencing the system’s decisions in unfamiliar situations.
This perspective supplements the traditional model of an agent as a single-objective maximizer. Real-world neural networks may not always behave like perfectly rational agents equipped with explicit utility functions. A model may contain multiple behavioral tendencies shaped by different parts of its training history, with different tendencies becoming active in different contexts.
Understanding this process of value formation may help researchers determine whether a system that behaves well during training will preserve the same preferences when deployed in unfamiliar environments. It may also help answer whether human-desired values can generalize reliably and under what conditions dangerous tendencies are likely to emerge.
Research Style and Influence
Turner’s published research reflects a distinctive ability to work across multiple levels of analysis. He studies both abstract questions in decision theory and concrete neural activity within real language models. He is concerned with the potentially extreme long-term risks posed by advanced AI, while also emphasizing mechanisms that can be tested experimentally today. His research trajectory is therefore not simply a transition from theory to application. Rather, it represents an ongoing effort to test theoretical judgments through model-based experiments and then use experimental findings to revise our understanding of intelligent agents.
Turner also places considerable importance on clear communication and open discussion. Through his personal website, he systematically organizes his research results, perspectives, and related essays, making questions that might otherwise remain confined to specialist papers accessible to a broader audience. His writing often emphasizes conceptual boundaries and carefully distinguishes among conclusions that have been formally established, phenomena observed through experiments, and hypotheses that still require further testing. This distinction is particularly important in AI safety, a field that combines mathematical proof, empirical research, and predictions about future systems. Without clearly separating different levels of evidence, possibilities can easily be presented as certainties.
Overall, Alex Turner represents a group of AI safety researchers who combine concern about long-term outcomes with a strong commitment to empirical investigation. His work on power-seeking has helped the research community discuss agent-related risks more precisely. His research into activation engineering and refusal mechanisms has introduced new technical approaches for understanding and controlling large language models. Shard Theory, meanwhile, has broadened how researchers think about the formation of machine values.
A consistent question runs through all of this work: Why does a powerful AI system take a particular action, can humans understand the reason before the action occurs, and can reliable methods of control be established?
No complete answer to this question yet exists. Turner’s contribution does not lie in claiming that AI alignment has already been solved. Rather, it lies in breaking broad concerns about AI safety into concrete problems that can be defined, formally analyzed, and experimentally tested. At a time when AI capabilities are advancing rapidly, research that combines awareness of long-term risks with rigorous technical investigation is an essential part of understanding the intelligent systems of the future.
References
- Alex Turner, “Research,” https://turntrout.com/research
- Alex Turner, “About,” https://turntrout.com/about
- https://turntrout.com/why-i-left-google-deepmind