Red Teaming o1 Part 2/2– Detecting Deception with Marius Hobbhahn of Apollo Research artwork

Red Teaming o1 Part 2/2– Detecting Deception with Marius Hobbhahn of Apollo Research

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

September 14, 2024

In this Emergency Pod of The Cognitive Revolution, Nathan provides crucial insights into OpenAI's new O1 and O1-mini reasoning models.
Speakers: Nathan Labenz, Marius Hobbhahn
**SPEAKER_2** (0:02)
Hey, everyone. I'm excited to share Turpentine's newest show, This Won't Last. The show includes Logan Bartlett of Redpoint, Keith Rabois at COSLA, Kevin Ryan at AlleyCorp, and Zach Weinberg, a biotech founder and investor. They're not here to promote a portfolio company or share marketing blurs. If you were curious about what VCs talk about in group chats, when the cameras are normally gone, they're releasing footage for a short time. A link is in the description. Listen wherever you get your podcast.

**Nathan Labenz** (0:32)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together, we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello, and welcome back to a special emergency pod edition of The Cognitive Revolution. As the entire AI world reacts to OpenAI's announcement and same-day release of their new O1 and O1 Mini reasoning models, I sought out members of the O1 Red Team to get their takes on the new models' capabilities and safety profile, as well as the current state of OpenAI's approach to pre-release safety testing. I'm really grateful that within just hours of my reaching out, I had the opportunity to speak with Marius Haban from Apollo Research, and Leonard Tang, Aidan Ewart, and Brian Huang from Hayes Labs. While these two conversations are certainly not all you need to understand the new models, I do believe they provide a valuable perspective. And I'm glad to say that recent drama surrounding OpenAI notwithstanding, it seems that they've done a pretty good job with the O1 testing and release process. While I would have ideally liked to see our guests granted a bit more time for open-ended exploration, they did have a few weeks to conduct automated testing, which considering that these are funded organizations with full-time teams dedicated to building test suites in advance of new model releases does seem rather reasonable. I was also particularly pleased by how candid they were able to be in these conversations, and especially with the fact that Apollo had the opportunity to contribute directly to the O1 system card in a way that they ultimately felt very good about. From everything we've learned, it appears that the O1 models were created by applying intensive reinforcement training to the GPT-40 class of models. Remembering that GPT-35, the RLHF version of GPT-3, was released roughly two years later than the original, I think it's reasonable to think about the O1 models as a sort of GPT-45.
Where GPT-4 class models were already closing in on expert level performance on many routine tasks, O1's reasoning abilities are now enough to match or even exceed expert performance in many areas, while also expanding the scope of problems they can solve to include those that require more task decomposition and planning, trial and error, and other familiar forms of reasoning. This is more or less what I expected OpenAI to release next, and I think the nature of this model helps contextualize a number of recent statements made publicly by or otherwise attributed in the press to leadership at OpenAI, Anthropic, DeepMind, and Microsoft. Capabilities have clearly not plateaued. It had just been a while since the last major data point. Recent efficiency gains have been amazing, but models that can reason at length could easily more than offset them, particularly if they drive another major increase in demand. And the sort of detailed reasoning and problem-solving traces that O1 can produce are exactly the sort of synthetic data points that could get us over any natural data wall as leading labs continue to scale. As such, it's no surprise that OpenAI is not sharing the full chain of thought with users, and it's easier all the time to understand how Anthropic might believe that leading developers in 2025 or 2026 could get so far ahead of the field that nobody else has a chance to catch up. Safety-wise, meanwhile, it again seems that model capabilities and alignment are mostly highly correlated. O1 is harder to jailbreak, largely because it reasons more effectively in general, and this includes reasoning about what it should and shouldn't do. For now, overall, it seems that we're still in the sweet spot, where the potential utility of AI systems is tremendous, but the risks of major harm remain relatively minimal. And yet, at the same time, there are reasons to doubt that this trend will continue all that much further into the future. Apollo's work demonstrates that these models are more capable of subtle deception than previous generations, and they also show signs of potentially dangerous properties, including instrumental convergence and power-seeking, which AI safety researchers have been warning us about for years now. Again, this is far from the last word on this subject. As always with language models, there are many unknowns. And with that in mind, I invite all of you to do your part in exploring and characterizing the many different aspects of these new models. A key question will be just how capable AI agents become. I'll be watching that closely, and I'll be open to changing my assessment as I learn more. I'll absolutely keep you updated if I do. As always, if you're finding value in the show, we'd appreciate it if you'd share it with friends, write a review on Apple Podcasts or Spotify, or leave us a comment on YouTube. And we always welcome your messages either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. Now, let's hear from two organizations that have tested the latest models more than anyone else outside of OpenAI. I hope you enjoy my conversations with Marius Hophan of Apollo Research, and Leonard Tang, Aidan Uart, and Brian Huang of Hayes Labs. Marius Hophan, founder and CEO of Apollo Research, welcome back to The Cognitive Revolution.

54 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000669529789