**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share a cross post from the 80,000 Hours Podcast, featuring a conversation between host Rob Wiblin and Buck Shlegeris, CEO of Redwood Research and a leader in the nascent but rapidly evolving field of AI control. For those unfamiliar with the term, AI control represents a qualitatively different approach to managing AI risks. Rather than trying to create perfectly trustworthy AI systems from the start, Buck and collaborators are developing frameworks that they hope will allow us to deploy and get useful work from AI systems even if they remain fundamentally misaligned. This strategy feels particularly timely and important considering the increasingly sophisticated bad behaviors we've seen from frontier AIs over the last six months. From the scheming behaviors documented in O1 by Apollo Research, to the alignment faking that Buck's colleague Ryan Greenblatt demonstrated in Clawed, to the dramatic rise in flagrant reward hacking behaviors we're now seeing in O3, which by the way, OpenAI's own research suggests won't be easy to train away, at least for as long as we remain in the reinforcement learning scale-up era. All in all, it seems we're headed for a world of increasingly powerful and autonomous AIs that we can't fully trust, and thus have no choice but to figure out how to control. With that in mind, what I particularly love about Buck and team's work is their first principles approach to understanding AI systems and their unique strengths and weaknesses on their own terms. Critically, and in sharp contrast to humans, AI systems operate in information environments that their developers fully control, and they can thus be monitored and tested in totally different ways. Buck shares a number of great ideas in this conversation, including always-on monitoring systems that use weaker models to efficiently review model outputs and then escalate potentially concerning behaviors to stronger models and eventually to humans when needed, resampling strategies that discard potentially bad behaviors and simply have the AI try again, thus also helping us identify areas where models are particularly misaligned, paraphrasing strategies that attempt to disrupt any attempts models might make to communicate with one another in undetected code, factored cognition architectures that give models only the minimum information needed to do useful work, hopefully thereby making it much harder for them to execute long-time horizon attacks, and honeypotting strategies that tempt models to do something bad and then allow us to systematically study their capabilities and tendencies after they've taken the bait. Of course, there are likely to be important practical tradeoffs between safety and performance and safety and cost, and as you'll hear, Buck is very focused on developing techniques that frontier model developers will hopefully find cheap and easy enough to implement.
This won't solve everything. I don't think anything will, but it seems complementary to other efforts like interpretability and alignment science, and I really do see huge value both in the incremental security that AI control systems will provide as part of a broader defense-in-depth strategy, and also in the credible concrete proof of malintent that they might uncover if it turns out that the AIs truly are out to get us. As always, if you're finding value in the show, I'd appreciate it if you'd take a moment to share it with friends, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we welcome your feedback as well, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. Finally, a quick reminder, I'll be speaking at the upcoming Imagine AI Live event, May 28-30 in Las Vegas, the Adapta Summit, August 12-13 in Sao Paulo, Brazil, and the Enterprise Tech Leadership Summit, September 23-25, again in Las Vegas. Tickets are on sale for each of these events now, and we have links in the show notes. If you'll be there, please do ping me so we can meet up in person. For now, I hope you enjoy what I think is a really crucial conversation on AI control, from the 80,000 Hours Podcast with host Rob Wiblin and Buck Shlegeris, CEO of Redwood Research.
**Buck Shlegeris** (4:02)
At Google, probably less than 1 percent of your employees are spies. When an employee proposes an action that's potentially sensitive, if you have a different employee review it, that other employee probably isn't also a spy. Whereas in the AI situation, it's very correlated, right? They're probably either non-scheming against you or all scheming against you. The simplest thing you do once you catch the AI trying to escape, is you try to vary the space of escape attempts, right? Let the AI think it succeeded, simulate what it would see, and look at what it does from there. Suppose there's some task you wish the AI would do for you. It's plausible that you can convince the AI to do the thing through this mechanism. I recommend against getting a lot of work out of AIs that believe they've escaped because it seems pretty bad vibes. Kind of an interesting thing about models that are like egregiously misaligned is that all they wanted was to take over. And so from their perspective, even if you did a great job of controlling them, they're glad to exist, right? Like they thank you for the gift of bringing them into existence instead of some different AIs. Five years ago, I thought of misalignment risk from AIs that were capable of obsoleteing AGI researchers as a really hard problem. Whereas now, to me, the situation feels a lot more like, man, we just really know like a list of 40 things, where if you did the 40 things, none of which seem like that hard, you'd probably be able to not have very much of your problem. But then I've just also updated drastically downward on how many things AI companies have the time slash appetite to do.
156 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000706240600