**Erik Torenberg** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share an episode of the Future of Life Institute podcast, to which I've been a long-time subscriber and where I've been twice honored to appear as a guest, featuring a conversation between Jeffrey Ladish, executive director of Palisade Research, and host Gus Dacher. This cross post came about as I was preparing to interview Jeffrey myself. I had reached out to Jeffrey after seeing Palisade's recent work on reward hacking by reasoning models, and even scheduled a time to record. But Gus beat me to it, and after listening to this conversation, I thought that I could save Jeffrey some valuable time by cross posting instead, and really appreciate Gus for allowing me to do that. Palisade Research studies dangerous capabilities of AI systems, particularly focusing on loss of control scenarios. And as you'll hear, Jeffrey, who previously helped build the information security program at Anthropic, is an AI industry insider who believes that we're rapidly approaching the time when AIs will be sufficiently capable of hacking, deception, and long-term planning so as to present clear and present dangers. He also reports that his friends working in research at Frontier Labs often say that while they're increasingly fearful of the overall trajectory of AI development, they ultimately feel that their hand is forced by competitive pressures to keep moving forward. In this conversation, Jeffrey describes two broad ways that humans could conceivably lose control.
Acute crises, in which superhuman AI systems actively work against human interests, and slower moving scenarios, where society gradually but irreversibly shifts more and more decision-making responsibility to AI systems. He also goes into detail about their recent research into reward hacking by reasoning models in the context of chess games. As we've seen repeatedly now, models trained with reinforcement learning are more prone to a variety of bad behaviors. And as a recent paper by OpenAI showed, this is not an easy problem to solve. Toward the end, Jeffrey outlines what he thinks we should do about all this, advocating for greater coordination among AI labs, more transparency about capabilities, and potentially restricting further development of the most dangerous capabilities while continuing beneficial research and deployment. Now, that probably won't happen, barring a sufficiently shocking and damaging incident, but research of the sort that Jeffrey and team are doing is becoming more important all the time. So I'll definitely be following their latest results and look forward to discussing in a future episode with Jeffrey as well. As always, if you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post a review on Apple Podcasts or Spotify, or share any feedback or topic and guest suggestions that you have, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. Now, I hope you enjoy this conversation about AI reward hacking research and the big picture of AI risk and strategy. With Jeffrey Ladish of Palisade Research from the Future of Life Institute Podcast.
**Gus Stoker** (3:00)
Welcome to the Future of Life Institute Podcast. My name is Gus Stoker, and I'm here with Jeffrey Ladish from Palisade Research. Jeffrey, welcome to the podcast.
**Jeffrey Ladish** (3:09)
Hey Gus, it's great to be here.
**Gus Stoker** (3:11)
Fantastic. Maybe start a bit by telling us about what it is you do at Palisade.
**Jeffrey Ladish** (3:17)
Yeah, happy to.
So we are trying to study risks from emerging AI systems, and in particular, we are trying to better understand loss of control risks. And so, you know, this both looks like trying to understand, you know, sort of what are some of the strategic capabilities that are emerging in AI systems? You know, where might they act out in ways that will be hard to control? And then we are trying to sort of present things that we think we know about this to the public, to policymakers, to help people better understand, what is this weird situation we're in? Like, what is happening right now? And can we make sense of it? And can we make sense of it as a society and make good decisions about better paths to AI development?
**Gus Stoker** (3:56)
Yeah, there are many specific examples of potentially dangerous capabilities that I want to dig into. And you have a bunch of awesome papers about those. But maybe let's start with the beginning. What's the situation that we are in right now?
**Jeffrey Ladish** (4:09)
Yeah, so I was at Anthropic a few years ago. And, you know, I had this moment where I first used Cloud. And this was before Chad GPT was released. And I was like, what is happening? Like, I had seen GPT-2, I had seen GPT-3. And I was like, okay, this is pretty impressive. But, you know, I don't know how smart it really is. I started talking to Cloud. And I had a skin infection in my arm. It was swelling up. And I started asking Cloud, like, do I need to go to the emergency room? And Cloud was just very helpful at being like, well, yep, if there's swelling, if there's redness, if there's these specific signs. And I was like, actually, I have all of those. So I went immediately to Urgent Care. And they're like, yeah, you need antibiotics right now. And I was like, oh, wow, that was way faster than my doctor. Okay, these things are actually smart. Okay, I need to reorient. And I think that was when I first realized that at a visceral level, that, oh my God, scaling works. You can just take GPT-2, you can just throw in more data and more compute, and you actually get out intelligence. And so I think where we're at right now is this playing out several years later, right? Which is, the AI systems are actually getting smart, the models want to learn, you throw in more data, more compute, there's various methods that you use.
86 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000701949466