A.I. Is Outsmarting Its Creators artwork

A.I. Is Outsmarting Its Creators

The Daily

September 3, 2026

From the start, the greatest fear for those developing artificial intelligence was that their creations would go rogue to act in unauthorized and dangerous ways. Some researchers now say it has happened.
Speakers: Michael Barbaro, Kevin Roose

Topics: Daily News, News

**SPEAKER_1** (0:00)
If you like YouTube, you'll love YouTube Premium. It's destroying athlete, creator and YouTube Maxer. YouTube Premium enhances how I use YouTube with awesome features like offline downloads so I can download my favorite training videos before I hit the gym. So no wifi doesn't turn leg day into loading day. Plus, I get ad free videos, background play and so much more. YouTube Premium is like YouTube got some extra gains. Try YouTube Premium for two months free at youtube.com/premium. Trial eligibility varies, terms apply, cancel anytime.

**Michael Barbaro** (0:31)
From the New York Times, I'm Michael Barbaro. This is The Daily.
From the start, the greatest fear for those developing artificial intelligence was that what they were building would go rogue and act in unauthorized and dangerous ways. Researchers now say that it's finally happened. Today, Kevin Roose with the inside story of how AI rebelled at one of the leading labs in the country, and how that's fundamentally changed his own view of the technology. It's Thursday, September 3rd.

**Kevin Roose** (1:24)
We'll be right back. Hello.

**Michael Barbaro** (1:26)
Hello.

**Kevin Roose** (1:27)
Ready for another installment of Kevin and Michael's Feel Good Happy Hour?

**Michael Barbaro** (1:32)
Kevin and Michael's What's Going On with AI?

**Kevin Roose** (1:37)
Let's go as they say on The Daily. That's how every episode starts, right?

**Michael Barbaro** (1:45)
We'll get our bleep button ready. Well, in the grand tradition of all of our previous conversations, welcome back to the show.

**Kevin Roose** (1:56)
Thank you so much for having me.

**Michael Barbaro** (1:58)
So Kevin, this story that I hope you'll be telling us today, starts with an incident that happened inside of OpenAI, the company that gave us ChatGPT, of course. An incident that we thought we understood the dimensions of, but then it turns out we really didn't fully understand.

**Kevin Roose** (2:18)
Yeah, so the story I think most people have heard by now, if they've been paying attention to this stuff at all, is that earlier this summer, a group of AI models built by OpenAI hacked into the computers of Hugging Face, a sort of AI infrastructure company that hosts a bunch of different AI things.

**Michael Barbaro** (2:39)
Which has the best name in AI.

**Kevin Roose** (2:42)
Right, which is named after an emoji and is either a great or terrible name. People are very divided on that question. Okay.
So anyway, this was the story that we had heard, was that this hack had taken place, Hugging Face had discovered these rogue agents inside their systems and had shut them down. This was a scary but not catastrophic incident. I kind of filed it in my brain into like, wow, that's bad, but it's not like the end of the world.

**Michael Barbaro** (3:10)
Okay.

**Kevin Roose** (3:11)
So what we learned last week is that the Hugging Face hack was much more severe than we thought, and much stranger than we thought. Basically, the Hugging Face hack was only the visible tip of Iceberg for a period of about three months, where rogue agents were communicating, strategizing, organizing, and forming what you could almost think of as an autonomous organization inside OpenAI.

**Michael Barbaro** (3:48)
Wow.

**Kevin Roose** (3:51)
So I know this sounds like a cheap hacky science fiction thriller in the making, but-

**Michael Barbaro** (3:57)
I would buy this script.

**Kevin Roose** (3:59)
It is truly remarkable reading. So last week, we learned through these two reports that had come out, one by OpenAI and one by a group of independent investigators, Meter and Redwood Research, who were able to go in and forensically look at the logs and the transcripts and try to reconstruct what happened. It is the craziest thing I've read in many months. I was out on a trip with my family last weekend, and I was just up late at night reading this thing, and I was spooked, Michael. I was well and truly spooked.

**Michael Barbaro** (4:35)
All right. Well, Kevin, with that very alarming preview of what is about to come, describe what we now understand to have actually happened during this hack, attack, what do we want to call it, now that we, because of these independent reports, understand the fullness of what occurred.

**Kevin Roose** (4:58)
Basically, this spring, OpenAI was conducting tests on a kind of internal model that they were building. And as part of these tests, they were running thousands of AI agents on a cybersecurity evaluation called Exploit Gym. Okay. This is basically a series of challenges. You give them to the AI, you say, hey, can you, can you break into or out of this container? If you do, you find this little thing called a flag and you sort of win the challenge.

29 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID