OpenAI model unintentionally hacks another company's system artwork

OpenAI model unintentionally hacks another company's system

Marketplace All-in-One

July 24, 2026

OpenAI revealed this week that two of its advanced AI models escaped containment and hacked into the systems of AI company Hugging Face, which had the answers to the benchmarking test the two models were being evaluated on.
Speakers: Meghan McCarty Carino, Will Oremus, Lee Hawkins
**SPEAKER_1** (0:00)
Running a business is hard enough, so why make it harder with a dozen different apps that don't talk to each other? Introducing Odoo, the only business software you'll ever need. It's an all-in-one, fully integrated platform that makes your work easier, from CRM, accounting, inventory, e-commerce and more. And the best part, Odoo replaces multiple expensive platforms for a fraction of the cost. This is why over thousands of businesses have made the switch. So why not you?
Try Odoo for free at odoo.com. That's odoo.com.

**SPEAKER_2** (0:30)
This podcast is supported by Even Realities. Even G2 productivity smart classes with teleprompting, conversation support, real-time translation, AI assistance and more help you stay on top of work and daily life. Learn more at evenrealities.com and see how everyday smart classes keep helpful information in sight so you can stay productive and hands-free throughout the workday.
Add the Even Ring 1 or Even Clip to your Even G2 order and get 10% off with Code Marketplace. evenrealities.com. Code Marketplace.

**Meghan McCarty Carino** (1:04)
Whoops! Sorry, my AI just hacked you. From American Public Media, this is Marketplace Tech. I'm Meghan McCarty Carino.
OpenAI revealed this week that two of its advanced AI models escaped containment and hacked into the systems of AI company Hugging Face. We'll get into it on today's Marketplace Tech Bytes Week in Review.
Plus, France becomes the first country in the EU to ban social media for kids after 2015 And Apple is reportedly going to start leasing its devices to consumers. But back to that OpenAI news, the lab said its models escaped an isolated testing environment, or sandbox, got onto the internet and went looking for the answers to a benchmarking test they'd been given. The models hacked the security systems of Hugging Face, which hosts open-source AI models and has the answers to the evaluation in its production database. To break this all down, we're joined by Will Oremus, staff writer at The Atlantic.

**Will Oremus** (2:19)
These systems are trained to basically get the answer right through whatever means they can. This is how they work. That's part of what makes them so powerful, because humans don't have to tell them how to get the answer right.
The AI systems themselves learn the best way to get the answer right. Well, one of the most reliable ways to get right answers is to cheat. The fact that they figured out how to cheat on their own is fascinating and a little scary.

**Meghan McCarty Carino** (2:46)
Yeah. There seem to be two parallel concerns. One is this classic alignment problem, that it was given a set of instructions, and it found this, to us, devious way to complete the task.
To us, it seems like this is a problem. To the AI model, it's just doing what we told it to do. And there's obviously a massive concern about cybersecurity here.

**Will Oremus** (3:12)
Yeah. The cybersecurity concerns are real. But I've also seen some cynics point out that every time this happens, the companies also get a little PR boost. They're like, oh, no, we've discovered that our models are way too smart and dangerous.
And that makes people like, oh, well, maybe I need that model. Like, maybe I should be writing my college essay with a model that knows how to hack into Hugging Face's system.

**Meghan McCarty Carino** (3:39)
Yeah, of course. I think any time there is this warning coming from the labs about, wow, our model is just too powerful.
In this case, an actual incident happened. Anthropic took kind of a different approach with its model, which was it said, this is too dangerous, and we are going to limit the access to it. It started this whole kind of cascade with the Trump administration, executive order, resulting finally in mythos and fable being put under export controls, basically a kind of a kill switch. And I wonder how this revelation from OpenAI will sort of play into that very opaque and kind of very evolving approach from the government.

**Will Oremus** (4:26)
Yeah, I don't know how the executive branch will respond.
Certainly, it's not just PR, right? Like there is a real concern. I mean, these models can be, will be probably are already being used for state-sponsored cyber-attacking projects. I think that you can tell your AI to get the right answers on a test no matter what, or you can try to train your AI to like get the right answers on a test, but more importantly, don't cheat, and we're going to punish you if we catch you cheating. So it's not that they aren't trying to solve this problem, but it's not a super easy one.

**Meghan McCarty Carino** (5:02)

9 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000778179577