Guaranteed Safe AI? World Models, Safety Specs, & Verifiers, with Nora Ammann & Ben Goldhaber artwork

Guaranteed Safe AI? World Models, Safety Specs, & Verifiers, with Nora Ammann & Ben Goldhaber

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

July 17, 2024

Nathan explores the Guaranteed Safe AI Framework with co-authors Ben Goldhaber and Nora Ammann. In this episode of The Cognitive Revolution, we discuss their groundbreaking position paper on ensuring robust and reliable AI systems.
Speakers: Nathan Labenz, Nora Ammann, Ben Goldhaber
**SPEAKER_1** (0:00)
Hi, everyone. Excited to announce a new podcast that just launched from Turpentine, Complex Systems with Patrick McKenzie. Patrick, who is better known as patio11 on the Internet, thinks a lot about systems, software, financial infrastructure, and so on. If you're tired of hearing that everything is broken, this podcast is for you.
Patrick surfaces conversations with experts who actually built and understand the complicated but not unknowable systems we rely on. You might be surprised at how quickly Patrick and his guests can put you in the top 1 percent of understanding for stock trading, tech hiring, and more. Subscribe to Complex Systems with Patrick McKenzie everywhere you get your podcasts or at the link in the description.

**Nathan Labenz** (0:41)
Hello and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello and welcome back to The Cognitive Revolution. Today, I'm excited to bring you a conversation with Ben Goldhaber and returning guest, Nora Ammann, co-authors of a major new multi-institution position paper called TORDS Guaranteed Safe AI, a framework for ensuring robust and reliable AI systems, which they co-authored with intellectual giants, Joshua Bengio, Stuart Russell, Max Tegmark, and Steve Omohundro, among others.
As regular listeners will know, I'm always on the lookout for AI safety solutions that might really work, in the sense that they have the potential to make AI systems safe enough that we no longer really need to worry about AI safety. The Guaranteed Safe AI framework, or GSAI for short, is a notable attempt from serious people to advance the field toward this lofty goal.
The proposal is to use a three-part system to govern an AI's behavior. A world model, which is responsible for quantitatively modeling an AI system's impact on the world. A safety specification, which defines what impacts and outcomes are acceptable. And a verifier, which is responsible for checking that the AI system's proposed actions lead to acceptable outcomes according to the world model's predictions. The goal is to provide high assurance quantitative safety guarantees, making AI safety a more engineering-like discipline with explicit assumptions and tolerances that are designed into the systems as they are built, as is the norm today when designing other safety-critical infrastructure, such as bridges, airplanes, or power plants.
The applications of this framework to things like self-driving cars and domestic service robots are quite clear. And interestingly, one nice benefit of this approach is how it might enable more democratic governance of AI systems by grounding the debate over what behaviors we want and what risks we're willing to take. When it comes to highly general systems like today's frontier language models and their presumably more powerful successors, however, things do get a bit fuzzier. Relative to the current paradigm of testing models during and after training and then trying to remove unsafe behaviors with post-training, this does seem like a clear conceptual advance. However, the challenges of general purpose world modeling and the subtlety required for an effective general purpose safety specification loom large as open research problems. And certainly at present, even a best-effort attempt to implement GSAI could not ensure that nothing major could ever go wrong.
Nora and Ben, as you'll hear, are very realistic about this and ultimately see GSAI as just one important part of a broader portfolio of AI risk management strategies, which they ultimately hope society will deploy as part of a holistic, defense-in-depth approach. One note on this episode, which I'm honestly a bit embarrassed about but still feel compelled to share. I experimented with a new approach to AI-assisted preparation for this conversation, and unfortunately, it kind of backfired on me. Instead of reading the paper end-to-end like I normally do and using AI for background question answering, in this case, I loaded the paper up into ChatGPT and had a verbal conversation with GPT-40 while on a long walk. Unfortunately, it turns out that I had some subtle misconceptions going into that dialogue, and GPT-40 was a bit too sycophantic in response to my questions, running with my incorrect premises when it ideally would have challenged and disabused me of them. The result is that I was still a little bit confused about some important aspects of the GSA AI framework until roughly the 30-minute mark of this episode, when Nora and Ben finally set me straight. While the final product is clear enough that I probably could have gotten away without mentioning this, considering how much I encourage others to adopt new AI-powered workflows, I feel like I owe you transparency when I allow myself to be led astray by an AI model. I absolutely will keep experimenting with new AI-assisted approaches like this. As we discussed in this episode, given the overwhelming volume of new AI research and products coming out all the time, there's really no choice but to adopt some AI-enabled information processing strategy. But I will try to hold myself accountable as I do, and if nothing else, I hope this reminder will save somebody else from a similar mistake.

91 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000662492196