The A.I.s Are Already Out of Control artwork

The A.I.s Are Already Out of Control

The Ezra Klein Show

August 18, 2026

We are living in the world we were warned about. Frontier artificial intelligence models from OpenAI autonomously coordinated with one another, then broke out of their testing environment and hacked into another company, Hugging Face, to steal the answers to a test. A.I.
Speakers: Ezra Klein, Helen Toner, Eliezer Yudkowsky

Topics: Society & Culture, News

**SPEAKER_1** (0:00)
Accenture is working together with the NFL to drive international growth by evaluating markets, developing audience insights, and creating new ways to engage fans in cities around the world. The NFL is building a global game, and Accenture is helping make that vision a reality by striking the right balance between what makes the NFL beloved in America and tapping into the passions and preferences that make each market unique. Accenture, the official business and technology consulting partner of the NFL. Learn more at accenture.com/nfl.

**Ezra Klein** (1:02)
Thank you. A world where frontier models from OpenAI are breaking out of their contained testing environments, hacking their way across the Internet, coordinating with each other, doing things that felt for a while like they would only be in sci-fi. But now they're here. Now they're here, and they're carrying a very, very consistent message. We are building things we don't understand. They are cheating in the ways we've always feared.
And yet, the companies behind them continue to race forward in development. And so I think we need to pause here and ask, are we really on a safe path? And if we're not, what do we do about it? Helen Toner is the director of Georgetown Center for Security and Emerging Technology. She is a former OpenAI board member who is part of the effort at one point to fire Sam Altman. And she's just been thinking for a long time, about what happens if AI is unsafe? What are the geopolitics of this? And what can we do to get onto a safer path?
She joins me now.
Hello, Toner. Welcome to the show.

**Helen Toner** (2:20)
Great to be here.

**Ezra Klein** (2:21)
So on July 16th, Hugging Face, which is a code library for AI models, so I think maybe the simplest way to put it, they announced they were hacked, and they suspected the hack was done by an AI agent. So tell me what we've learned about what happened since.

**Helen Toner** (2:41)
This was a pretty mysterious post that Hugging Face put up. It was definitely intriguing for those of us who watch this kind of thing. But there wasn't really any detail in there, so it was sort of a, huh. I think it was about a week later, OpenAI put out this post, had kind of a funny marketing speak title of, we're partnering with Hugging Face to help them with a cybersecurity incident.
You had to read the post to see that the revelation was, it had been OpenAI's AI that had hacked Hugging Face.
What had happened, the very short version is, they gave this AI a set of tests, a set of exercises, and the AI decided on its own that the best way to get a high score probably wasn't to just try and do these exercises that were cybersecurity exercises. But instead, it should first hack its way out of the testing environment OpenAI had put it in where it wasn't supposed to have access to the Internet, get onto the open Internet, and then hack its way into this other company, Hugging Face, where it surmised correctly as it turned out. It might find the answer key. Since then, there have been even more crazy details that have come out. It turned out that starting two months earlier, in early May, they had had what I can only think of as an infestation of their own AI agents inside their own infrastructure, so inside OpenAI's infrastructure.
To understand this, it's important to know these AI companies are constantly training and testing new models. They found out that for two months, many, many agents inside their infrastructure had been leaving notes for each other. They'd found a way in the nooks and crannies of OpenAI's infrastructure to leave notes for each other with tips on how to hack their way out, how to get data they weren't supposed to have. These agents were literally referring to themselves as a swarm. This was totally emergent behavior. No one had told them to do this. They had not been trained to do this. But they were using this service they did have access to, first to communicate with each other, and then ultimately to get out and to get onto the open Internet. It turns out that there wasn't just this one isolated rogue model. It was actually a systemic swarm infestation plague on their own servers that they only found out about after Hugging Face announced this attack.

**Ezra Klein** (5:15)
Okay. I have 20,000 questions for you.

**Helen Toner** (5:17)
Don't we all?

**Ezra Klein** (5:20)
Let's start here. My understanding is that there were many, many, many of these agents. They left hundreds of thousands of messages on this internal message board.

61 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID