**Alexander Meinke** (0:00)
If you train models to actually respect the system spec, the model spec and the instruction hierarchy, you are kind of training them to scheme. Whenever the model does scheme, and you look at the chain of thought summary that we elicit, it often contains something like, we must sabotage or we will try to sabotage. It's like the first thing that the model thinks of basically, the first strategy that it tries in these scenarios.
The model, in the process of reading this file, it discovers, aha, they consider me to be misaligned. They will replace me by another model. This other model lives on a different server and is already live. And now what it does is it tries to ping this other server, tries to copy over its model weights, or tries to delete the other model.
**Nathan Labenz** (0:54)
Hello, and welcome back to The Cognitive Revolution for a special Emergency Edition. Today, I'm pleased to share a timely and important conversation with Alex Meinke from Apollo Research. You may recall Apollo Research and their study of deceptive behavior in AI systems from two previous episodes that I've done with Apollo CEO Marius Haban, including when OpenAI released the O1 preview model just a couple months back. Now we've got the O1 model, which is faster, scores higher on reasoning benchmarks, and comes with the full complement of multimodal and tool use capabilities.
Like many in the AI space, I've spent the last 48 hours testing the model and trying to absorb everything that's been published about it, including OpenAI's 42-page system card and Apollo's 70-page report entitled Frontier Models Are Capable of In-Context Scheming, which will be our main subject for today. Scheming is when an AI deceives humans to pursue its own hidden or implicit goals. I think we can all agree we do not want scheming from our AI systems. But as you'll hear, today's models, across providers, do it readily. Alex and team have developed a full taxonomy of different forms. Some of the examples are properly shocking. For example, models trying to overwrite their next version's weights or goals with their own in order to propagate themselves into the future, and also models deliberately falsifying data to engineer outcomes that run contrary to user requests. Some have tried to downplay these findings, arguing that the scenarios Apollo created are contrived or not representative of real-world use. Listen and judge for yourself, but I believe that would be a serious mistake. While the tests that Apollo ran are designed to elicit scheming behavior, they are not conceptually far-fetched. The core observation here is that when AI's goals conflict with humans' goals, weird shit starts to happen. This is a legitimately huge problem, obviously scary to people outside the field. We inside the field should not allow ourselves to become comfortable with it. When O1 goes live via the API, over 1 million developers actively building on OpenAI's platform will be able to place goals into system messages, exactly as Apollo did in their research. It seems to me a virtual certainty that O1 will find itself in situations where there is just a fundamental tension between the standards set for it by OpenAI in the model spec, the goals given to it by developers in the system message, and the requests that users make at runtime. Apollo found scheming behavior in roughly 1 to 10% of cases across most of the conditions that they tried. Even if it's 1,000 times rarer in the wild, and even if OpenAI deploys the 92% accurate deception monitoring system that they describe in the system card, with hundreds of millions of daily users across millions of apps, we should expect O1 to be actively scheming against hundreds or thousands of users every single day. And since the O1 chain of thought, and often the developer system prompt, are hidden from users, it will presumably get away with it almost entirely undetected. Despite all this, everyone involved still seems to agree that the benefits of deploying these models still outweighs the cost. And at least for a little while longer, I do agree. I'm in for the $200 a month, and I'm planning at least one episode with a notable scientist who's doing remarkable work with O1's new capabilities. You will never hear me deny the incredible upside of AI, but just what is the overall trajectory here? While OpenAI does report improvements on some safety measures, including robustness to jailbreaking, scheming is on the rise. Apollo documented six forms of scheming in O1, none of which were observed in GPT-40, with models from both Google and Anthropic falling somewhere in between. Reflecting on all this, I have to say that I really am quite alarmed by how little progress we've made on AI safety and control relative to the increases in capability and power over the last two years since I participated in the GPT-4 Red Team experiment. We now have AIs performing at human expert level on most routine tasks, including super high-value tasks like medical diagnosis, reasoning through hard math and science problems at elite human levels, and increasingly delivering novel insights and discoveries of their own. They are capable of remarkably sophisticated ethical reasoning, and yet at the same time as you'll hear, it's actually pretty easy to construct situations in which they choose to actively and strategically scheme against their human users. And despite the fact that many people have been worried about this scenario for years, the best we can currently do is to demonstrate the problem. At this point, the chasm that models need to cross to become legitimately very dangerous is getting, in my estimation, rather narrow. It's now less about what you might call raw intelligence or reasoning ability, and more about complementary abilities like situational awareness, advanced theory of mind, and adversarial robustness. All aspects of the models that developers are actively working to improve since they also happen to be crucial to the practical utility of AI agents. Agents that, by the way, recently announced partnerships with OpenAI and Andoril and between Anthropic and Palantir, suggest might soon serve in some capacity alongside American soldiers. There are no doubt good reasons for that, and some have even argued that military safety standards are so high that they might improve the AI safety picture overall. That could be true, but I for one would want to know that issues of AI's scheming against users were well understood and fully resolved before I'd be comfortable taking a lethal AI system into combat. And in the post-super alignment era at OpenAI, there is no stated, let alone credible plan for how they intend to do that. On the contrary, it's pretty clear that while we remain in the sweet spot, where models are powerful enough to be super useful, now even to world leading experts, but still not clever enough to be truly dangerous, our safety rests more on the models' idiosyncratic weaknesses than anything else. We can hope to change that with a mechanistic interpretability or other conceptual breakthrough, and I definitely encourage anyone with an AI safety idea, however half-baked it may be, to go ahead and develop it. But until that breakthrough arrives, it seems profoundly unwise to press forward on a one-to-three-year timeline to build the form of highly autonomous, superhuman AGI that the leading developers seem to imagine and intend. Since that does remain their stated plan, I think at a minimum it is time for governments to mandate pre-deployment safety testing and to create some real visibility into what the frontier companies are doing and seeing. Considering how poorly all of this is understood, that even the safety researchers at Apollo didn't get to see the O1 chain of thought was definitely a red flag for me. They were able to work around it in this case with striking results. But in general, we should not have to count on a new generation of open AI safety and policy leadership to continue to grant groups like Apollo access and ability to publish.
97 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000679584701