**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, my guest is Jeffrey Ladish, Executive Director of Palisade Research, which studies the capabilities and motivations of today's AIs as part of its effort to better understand the risk that humans could irrevocably lose control of AI systems. We begin with Palisade's work on shutdown resistance, which showed that in both digital and physical environments, even when they're explicitly instructed to allow themselves to be shut down, LLMs sometimes take extraordinary actions, such as disabling the shutdown mechanism in order to extend their sessions and continue to pursue their goals. We get Jeffrey's take on criticisms of the specific techniques used in this research, his current understanding of why it is that models act this way, which he attributes not to a proper survival drive per se, but a strong task completion drive, and his perspective on the current state of alignment writ large. In short, while he does recognize that current models are aligned enough to be super useful and he does use them actively, he's not optimistic that current techniques will be enough to keep models in the so-called benevolent basin as frontier training methods shift toward longer and longer time horizon tasks and potentially multi-agent competitive environments in which deception would often be naturally rewarded, just as it is in nature itself.
From there, we turn to Palisade's latest work, in which they demonstrate that even recent open-source models, while not yet able to find zero-day exploits like Mythos can, are now capable of self-replication by repeatedly exploiting known cybersecurity vulnerabilities in order to gain control of new servers, setting themselves up to run on these new environments, and prompting their copies to continue doing the same thing. In light of these issues, I was keen to get Jeffrey's cybersecurity advice for AI agent users like me. He recommended that I think hard about the so-called lethal trifecta of giving your AI agent access to sensitive private information, access to previously unseen and untrusted content which could contain prompt injection attacks, and the ability to communicate externally. And I certainly will be. More importantly, he also offers his analysis of where things are going from here. He explains what the world looks like to an AI agent, handicaps the difficulty that they'll face in colonizing different environments, from personal laptops to hyperscalar data centers, and reminds us that even if cyber defenders gain a technical advantage in light of superior computing resources and early access to the best models, humans will remain vulnerable to social engineering and will likely end up being the weak link in the chain.
At the very end, I asked Jeffrey what technical solutions he finds most promising, and as often happens when I pose such a question to somebody who's been grappling with these issues for years, he expressed enthusiasm for multiple lines of work, from compute governance to interpretability based monitoring, but ultimately concluded that the only strategy he really believes in is an international agreement to refrain from using recursive self-improvement to trigger an intelligence explosion. At least until we have a much better understanding of how to design and control AI motivations.
Overall, it's an arresting picture, but I hope you enjoy this mind-expanding look at what AI systems can already do today, and what it might look like for humanity to begin to lose control, with Jeffrey Ladish of Palisade Research. Jeffrey Ladish, Founder and Executive Director at Palisade Research. Welcome to The Cognitive Revolution.
**Jeffrey Ladish** (3:37)
Thanks for having me.
**Nathan Labenz** (3:38)
This has been a long time coming. We've met a few times at different events over the years, and I cross-posted an episode that you did on another podcast some time ago. And I'm glad to finally be doing one of these live. So it should be a very interesting conversation, because you are right in the thick of it right now, at the heart of where AI capabilities are going vertical, and the consequences are going from theoretical to practical concern, even for obscure folks like me on a day-to-day basis, in a pretty compressed time frame. So I'm going to be interested to hear both in the weeds details about the research that you've been recently doing and the observations that you guys have made at Palisade, and then really also looking forward to a broadened out conversation on what can I do about this, if anything, to protect myself, and what does it mean as we go forward into the very foggy AI future?
**Jeffrey Ladish** (4:34)
Yeah.
Well, Nathan, I remember a year and a half ago, my team and I went to DC, and the thing we were doing in DC was we were briefing a lot of folks in Congress, in the admin, a lot of different staff, and we had a presentation that was like, hey, autonomous cyber agents are coming. Like AI agents that can hack pretty autonomously at scale are on the way. And the reason we know this is because OpenAI just released a model called O1 that has been trained via reinforcement learning on actual programming problems. And the scale up from O1 to O3 is incredible. And I don't know if O3 was even out yet. But like in that time when we had moved from just like a pre-training regime, where it was just like throwing a bunch of human data to the point where, no, we can actually train these models by, they can do trial and error on their own. They can do exploration on their own. We have to build reinforcement learning environments for them.
116 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000769354640