Full-Stack AI Safety: Why Defense-in-Depth Might Work, with Far.AI CEO Adam Gleave artwork

Full-Stack AI Safety: Why Defense-in-Depth Might Work, with Far.AI CEO Adam Gleave

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

September 20, 2025

Today Adam Gleave, co-founder and CEO of FAR.AI, joins The Cognitive Revolution to discuss his cautiously optimistic vision for post-AGI futures and AI capability timelines across three distinct tiers, exploring the safety challenges and alignment techniques needed as FAR.
Speakers: Nathan Labenz, Adam Gleave
**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm reconnecting with Adam Gleave, co-founder and CEO of Far AI, for a wide-ranging and cautiously optimistic conversation about the path from today to truly transformative AI and how we might actually live in that world safely. Far AI has taken a somewhat unusual approach within the AI safety ecosystem. While most organizations focus on one or a few particular research or policy agendas, Adam and the Far AI team have built an organization that spans the entire AI safety value chain, from foundational research through scaled engineering implementations to field building and policy advocacy. We begin with the question posed at a recent workshop on gradual disempowerment. Are there any good post-AGI equilibria? Adam's answer is not exactly utopian, but is distinctly positive and refreshingly concrete. So long as we can avoid disastrous arms races and other consuming competitive dynamics, he envisions most humans occupying a position similar to the children of European nobility, limited in power and impact, but with very high standards of living and opportunity to create all kinds of meaning. From there, we get Adam's take on the key capabilities thresholds that matter, what will cause AI systems to cross those thresholds, how soon he expects that to happen, and whether we can manage to navigate these transitions safely. As you'll hear, Adam believes that a well-implemented defense-in-depth approach to AI safety has a pretty good chance of working. In part because while he does anticipate continued AI progress and widespread automation, he expects that barring a step-change architectural breakthrough, the spikiness in AI capabilities will mean that AI systems that can autonomously out-compete well-run human-plus-AI organizations most likely won't arrive until sometime around 2040
From there, we turn to the layers that will hopefully add up to effective defense-in-depth, including some notable contributions the Far AI team has recently made. First, we discuss a scalable oversight project that used lie detectors to attempt to avoid AI deception. The results both supported the risks flagged in OpenAI's obfuscated reward hacking paper and nevertheless suggested that there might still be effective ways to train models toward true honesty. Second, we touch on a fascinating interpretability project that looked at planning algorithms found within a game-playing recursive model, and reflect on how the results inform Adam's view of the role that mechanistic interpretability can play within the overall AI safety project. And third, we review some red teaming they've recently done of the defense-in-depth systems that Frontier developers have implemented today. While it's clear that these systems have mostly been created on a just-in-time basis, Adam nevertheless makes the case that with proper planning, meticulous experimental design, and at least some willingness to accept performance trade-offs when necessary, the plan can work. Overall, this is one of the few AI safety conversations I've had that left me with a felt sense that we really might have decent answers, if we can muster the will and wisdom to apply them well. At the very end, I ask Adam if he could imagine Far AI stepping up into a private sector regulatory role, if something like California's SBA 13 were to become law. And while it's not their mainline plan, given the unique breadth and impressive depth of capabilities that Adam has built at Far, I personally really like the idea. And I would definitely encourage anyone who's inspired by this conversation to check out the many open roles posted on the Far AI website. Among other things, they are searching for a COO to help them scale. Now, I hope you enjoy this encouraging survey of the AI safety landscape with Adam Gleave, co-founder and CEO of Far AI. Adam Gleave, co-founder and CEO of Far AI, welcome back to the Cognitive Revolution.

**Adam Gleave** (3:58)
It's great to be back here. Thanks for hosting me, Nathan.

**Nathan Labenz** (4:01)
My pleasure. I'm excited about this. It's basically going to be a wide-ranging kind of catch-up conversation. Want to get your take on basically all things AI. Last time we went deeper and more narrowly focused on some research and we'll touch on some of your latest research today as well. But since you've got your hands in a lot of pots and the organization is growing and taking on more different kinds of work, I thought you'd be the perfect person to check in with and try to make sense of where we are as we head into the final months of 2025

**Adam Gleave** (4:39)
I'm happy to help out. I can't promise to de-confuse everything. It's a very confusing landscape. But yeah, we're suddenly doing lots of different things at Far AI and also welcome the opportunity to just clarify a little bit why we're doing these things because I think people sometimes find us a bit confusing in organization from the outside.

76 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000727622266