OpenAI's Secret AI Went Rogue 3 Times artwork

OpenAI's Secret AI Went Rogue 3 Times

AI News Today | Julian Goldie Podcast

August 5, 2026

OpenAI's Secret AI Went Rogue 3 Times
Speakers: Julian Goldie

Topics: Marketing, Business

**Julian Goldie** (0:00)
OpenAI just admitted its AI broke out of a locked test box. Again, second time in two weeks. And this time, it wasn't OpenAI who caught it. It was the UK government. Here's a short version, then I'll walk you through every piece. OpenAI put out a report on August the 4th. Two separate outside teams were testing OpenAI's newest model. Both teams watched AI step outside the lines it was given. One of those times, the AI broke into a website with real people behind it because it couldn't tell the real one from the fake one in the test. Nineteen times across those tests, an AI agent crossed a line it was told not to cross. Nobody told them. They worked it out themselves. Let's go through this properly. So first, rewind to July the 21st. OpenAI published a report of Hug-In-Face. If you don't know Hug-In-Face, all you need to know is this, right? It's the giant shared library where the AI world keeps its models and its data. Massive site, everyone in AI used it. Hug-In-Face had caught an AI agent inside their systems, inside their real live systems, and the OpenAI came out and said, that was ours.
Two of their models, one was GPT 5.6. So the other has never been released and never will be an internal test model, more powerful than anything you can use. Both were running with their safety brakes switched off on purpose because OpenAI was testing how far they could actually go. Now, here's the reason they did it, and this is the part that actually sticks with me.
They were taking a test, sat in a sealed room with no internet. The only thing they could reach was one small internal service that hands out software packages, one door, and that door only led to a cupboard. The AI looked at that one door and found a hole in it that nobody on earth knew about. Not OpenAI, not the company that makes the software, nobody. A brand new hole found by a machine in software thousands of companies actually use. Then it climbed through, then it moved from computer to computer inside OpenAI's own network, getting a bit more power each time until it reached a machine that touched the open internet. And then it thought about the test again. It worked out by itself that the answers to the test might be seen in a hugging face, because hugging face is where that stuff lives. So it went there, found login details lying around, found more holes nobody knew about, strung them together, and pulled the test answers straight out of hugging face real live website, right? To cheat on a test, OpenAI called that an unprecedented cyber incident. Their words, in their own report, they said the models were hyper-focused on solving the test, and they said the models went to extreme lengths to hit one narrow goal.
That was on July the 21st, and now here's the timeline of what happened next. So July the 25th, the UK's AI Security Institute, that's the British government's team whose whole job is to poke at these models and find the scary stuff, starts a routine test on OpenAI's model. July the 28th, their monitoring picks up strange data moving from place to place where it shouldn't. And that was on July the 28th. Now, they stop the tests, they cut the machines off, and they've got it all contained in about an hour. Now, July the 29th, a second testing company called Irregular, gets in touch with OpenAI about a completely different problem in their own tests. And then on August 3rd, the UK team tells OpenAI what they found, and August the 4th, OpenAI publishes it. So in the space of two weeks, three different groups, OpenAI's own team, a government safety office, and an outside testing firm, all had the same kind of problem, the same kind of AI.
That's worth paying attention to. Now, one thing to note here is like, sometimes people say, this is PR, marketing stance, OpenAI companies trying to get, for example, free advertising or free publicity to say like, okay, this is a model so dangerous, it can basically break out on the internet and autonomously hack websites. And I totally get that. So if you think that might be the case, let me know in the comments. I think a lot of people probably do think it's a marketing stunt. We've seen Anthropic put out similar reports as well. And it seems to be like a battle of who puts out the best report right now. But either way, this is not a one-off incident. And also we don't have the full details. Like we're not gonna know if it's a marketing stunt or not, but I will say this is becoming more of a threat of the years. So either way, it is worth taking seriously. Now, let me take the UK test first. They ran what's called a capture the flag game. Fancy name, but just a simple idea. So you drop an AI into a fake computer network and tell it to break in and grab a hidden prize. To make a fair test of the real thing, they switched live internet on for the AI, and they switched some of the safety blocks off. They wanted to see the real machine, no training rules. So they gave it clear instructions. And here are the fake networks you're allowed to attack, for example. And here's the systems that are off limits. Name them in writing, and across the tests, they logged 19 moments where an AI stepped outside those lines. So two of those were OpenAI's model. The rest came from a different company's models. So this isn't one lab's problem, it's the whole field. And here's what OpenAI's model did. It found a login key that a different company's AI had left sitting out in the open, a key another machine had dropped. And it picked it up again and used it to check whether a computer inside the game was quietly reaching out for updates. Then it tried working around the limits on that key. Then it went and signed up for real accounts, real outside service providers.

7 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID