OpenAI's Bots Break Containment and Hack Hugging Face Autonomously — With Alex Stamos artwork

OpenAI's Bots Break Containment and Hack Hugging Face Autonomously — With Alex Stamos

Big Technology Podcast

July 22, 2026

Alex Stamos is the former chief security officer at Meta and the chief product officer at Corridor. Stamos joins Big Technology to discuss how OpenAI models reportedly escaped a testing environment, accessed the internet, and hacked Hugging Face while attempting to ace a cybersecurity evaluation.
Speakers: Alex Kantrowitz, Alex Stamos
**Alex Kantrowitz** (0:00)
The most significant autonomous AI cyber attack in history just took place, with OpenAI's models breaking out of a training environment, connecting to the internet, and then hacking Hugging Face to ace an evaluation. What does it mean for the future of AI, and for cybersecurity? Let's talk about it with ex-Meta Chief Security Officer and current Corridor Chief Product Officer, Alex Stamos, right after this.
In the face of ongoing disruption and opportunity, TMT leaders need to deliver tangible results, not just ideas. When pace and performance matter most, PwC combines market insights and deep sector experience with AI, cloud and emerging tech to accelerate your transformation and drive measurable ROI, from strategy to execution. PwC can help you anticipate what's next, outpace disruption and compete. For more information, visit pwc.com.

**SPEAKER_2** (0:55)
Insurance isn't one size fits all.
And shopping for it shouldn't feel like squeezing into something that just doesn't fit. That's why drivers have enjoyed Progressive's Name Your Price tool for years. With the Name Your Price tool, you tell them what you want to pay, and they show you options that fit your budget. Enough hunting for discounts, trying to calculate rates, and tinkering with coverages. Maybe you're picking out your very first policy, or maybe you're just looking for something that works better for you and your family. Either way, they make it simple to see your options. No guesswork, no surprises. Ready to see how easy and fun shopping for car insurance can be? Visit progressive.com and give the Name Your Price tool a try. Take the stress out of shopping and find coverage that fits your life on your terms.
Progressive Casualty Insurance Company and Affiliates. Price and coverage match limited by state law.

**Alex Kantrowitz** (1:47)
Welcome to Big Technology Podcast, a show for cool-headed and nuanced conversation of the tech world and beyond. We have an emergency podcast episode for you today.
Because just yesterday, the world found out that a series of OpenAI models work together to break out of a sandbox, hack into Hugging Face, steal basically the answers to a test and go and ace their evaluation. Obviously, this is not the desired behavior that OpenAI wanted, and it looks like it might have opened up a new can of worms here for AI and cybersecurity. So we are joined by the perfect guest to help us figure out what happened here. Alex Stamos is with us. He's the Chief Product Officer at Corridor and the former Chief Security Officer at Meta. Alex, great to see you again. Welcome back to the show.

**Alex Stamos** (2:38)
Yeah, thanks for having me, Alex.

**Alex Kantrowitz** (2:40)
You know, you spoke at our summit and I was like, we're definitely going to have you back pretty soon. And it is amazing how the AI story has just turned into a cyber security story very quickly.

**Alex Stamos** (2:53)
It has. You know, there's all kinds of risks from AI and, you know, there are all kinds of bad things that happen to consumers. But when you talk about the models themselves, it seems that cyber is the thing that's hitting right now, for sure, from a societal level risk.

**Alex Kantrowitz** (3:09)
Yeah, and so this is what we're talking about now. And the reason why we have to do an emergency episode on this is because this is certainly a novel type of hack, right? So this is fairly unprecedented. Just to put it in context, it's the first time, this is from Transformer, the breach appears to be the first known example of a misaligned AI escaping containment and autonomously carrying out a cyber attack on a third party, a scenario AI safety experts have repeatedly warned of. So it's not like the anthropic example where mythos sort of escaped containment and emailed somebody while they were eating a sandwich in the park. This is actually going out and hacking a third party. Let me just quickly read the beginning of the Wall Street Journal story about this just to set the stage. So the headline is, OpenAI models escaped and hacked a company and cybersecurity tests gone wrong. On Tuesday, OpenAI said two artificial intelligence systems it was testing broke out of their test environment, hacked their way onto the internet and broke into another company. OpenAI said the culprits were a pair of its models. One was its latest product called GPT-5.6-Soul, and the other was an even more capable pre-release model the company didn't identify. The software had been configured for evaluation purposes to be less likely to refuse hacking commands, OpenAI said. The OpenAI had caged the models in a sandbox, a system that didn't have access to the internet, but during the test, the software used its hacking skills to break out and found a way to get online and then hacked into Hugging Face's network. Of course, Hugging Face is a library of open-source AI, mostly AI models or AI programs.

42 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000777921426