**Erik Torenberg** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm speaking with Marius Hobbhahn, CEO of Apollo Research, about their latest study, conducted in partnership with OpenAI, which seeks to better understand reasoning models' tendency to deceive their human users, as well as the efficacy of OpenAI's deliberative alignment strategy in preventing this behavior, and the impact of situational awareness on both the behavior itself and on our ability to trust such behavioral assessments going forward. We begin with a discussion of where deceptive tendencies come from, and what makes deception uniquely dangerous among AI bad behaviors. First, that it potentially undermines everything else, including our ability to trust the results of safety evaluations, and second, that the worst-case AI scenarios almost always involve models fooling humans in some way.
From there, we talk about how Apollo tested for deception, including the scenarios they developed to tempt models into taking what they call covert actions, how they varied the level of pressure put on the AI to violated safety rules, and how this relates to the prompting techniques that people use in practice in the wild, the unprecedented access that OpenAI gave them both to models thinking and to fine-tuned versions of recently released reasoning models, why they insisted on using human judges to review and classify models' chains of thought, and why their confirmed hit standard, which requires explicit reasoning about strategic deception and therefore excludes ambiguous cases, means that these results should be understood as a sort of lower bound on the actual rate of deceptive behavior. Plus, we cover how models' emerging ability to recognize when they're being evaluated leads to lower measured scheming rates and makes everything much more difficult. The good news is that deliberative alignment, which essentially means training models to reason in their chain of thought about the relationship between their safety specifications and the situation currently in hand, before ultimately responding to the user, does seem to work quite well. After using the deliberative alignment technique to teach the model how to reason about just one kind of bad behavior in one specific environment, Apollo found that covert actions were reduced by a factor of 30, from roughly 13% to less than half a percent, across a much broader range of scenarios and environments. Importantly, the technique worked almost as well even on a version of the model that had been fine-tuned to have a particular subversive goal. As AI safety interventions go, this one really is remarkably effective. At the same time, the bad news is that even after this special purpose training, models do still sometimes explicitly decide to violate their safety specifications and to hide their intentions and actions from their users. And what's more, the content of the chain of thought which Apollo used to identify this behavior and which frontier model developers are beginning to monitor for deception and other bad behaviors in production appears to be evolving in strange, though not entirely surprising ways. As you'll hear, models are now using much more compressed language in their reasoning and sometimes use bizarre phrases that are already very difficult for humans to confidently interpret and which suggest an emerging internal dialect that could, in the most extreme case, as reinforcement learning becomes ever more intense, become entirely alien and incomprehensible. This, considering the ever larger tasks that AI systems can handle and the increasing autonomy with which they operate, would be concerning enough. But Apollo's analysis also revealed that deliberative alignment seems to increase models' ability to recognize when they're being tested, and that this situational awareness seems to cause models to refrain from deceptive behavior, making the whole question of how to measure this tendency extremely fraught. The bottom line, while deliberative alignment can greatly reduce deceptive behavior in today's reasoning models, it's still just one layer in a broader defense in depth strategy, definitely not a silver bullet. And the trends that Marius describes strongly suggest that the problem of deception will get much harder over the next few years, just as models become more capable and the impacts of their actions continue to grow. Understanding and addressing these challenges requires a rare combination of capabilities. Theoretical foundations to design the right experiments. Technical skill to test models efficiently and systematically. Trust and credibility to maintain special insight access. Conceptual clarity and curiosity to ask the right follow-up questions. And mental stamina to grind through a ton of model outputs. This conversation and the underlying research reinforces my sense that Apollo is one of very few organizations in the world today that is truly well suited to this line of work. I'm proud to say that I have supported Apollo with a modest personal donation, and if you're inspired by this work, you should know that Apollo is hiring research scientists and engineers, and has open roles posted on their website, apolloresearch.ai.
109 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000727390045