HackGPT: How AI escaped the lab and went rogue artwork

HackGPT: How AI escaped the lab and went rogue

The News Agents

July 22, 2026

Have we just crossed the AI Rubicon? OpenAI has revealed that one of its models was involved in an "unprecedented cyber incident". One of its AI programs escaped a controlled testing environment, got onto the internet, and hacked into another AI start up without being asked to do so.
Speakers: Jon, Will Guyatt, Lewis, Andy Burnham
**Jon** (0:02)
This is a global production. OpenAI's own program has broken out of a security test and hacked a rival AI company, and no one told it to.

**Will Guyatt** (0:12)
This is exactly what happened in 1991's Terminator 2 judgment day.

**Jon** (0:17)
People have posited this as the nightmare scenario of AI where it's out of control.

**Will Guyatt** (0:24)
Google themselves six months ago admitted to shutting down an experiment because servers began speaking to each other and communicating in a language that the human users couldn't understand. The scariest thing about this is the next big accident will happen, and we will only find out about it.

**Jon** (0:40)
OpenAI has admitted an autonomous agent powered by its most advanced artificial intelligence models broke containment lines and hacked another company using stolen login details. For years now, we've been speculating, talking about perhaps dreading the idea that the AI systems that we're creating, those remarkable technological marvels, may start to not just do what we tell them, but to think for themselves. It's just possible we may have seen the strongest sign of that yet. Which begs the question, does AI have a mind of its own and has it started hacking us?
Welcome to The News Agents.

**Lewis** (1:29)
The News Agents.

**Jon** (1:30)
It's Jon. It's Lewis. And I want to talk about the Gothic novel. Oh yeah. And Frankenstein and Mary Shelley. Oh, I thought you were going to say Dracula. No. That was before I got my teeth done. When the monster takes control and has a mind of its own and takes revenge on its creators.
And are we seeing that today? That's my simplistic explanation. That's my easy parallel. A nice contemporary reference, then.

**Lewis** (1:53)
Yeah, that's what I am.

**Jon** (1:54)
A contemporary story.

**Lewis** (1:55)
Of course.

**Jon** (1:55)
OK, well, you stick to your Shelley and I'll just give a very brief description, Prometheus, as to exactly kind of what has happened before we talk to someone who really knows what they're talking about in terms of the tech. So this is a story which has emerged in the last kind of like 24 hours or so. Basically, it's a story which began during a routine safety test inside OpenAI, Sam Altman's company doing so much of the running on artificial intelligence. Engineers were basically running an experiment, seeing how two of its most advanced AI systems not available to the public would perform in a cyber security challenge. It was basically challenging this program or this set of programs to hack, and to basically properly test their limits, what they did was switched off the usual safety features that they have and restrictions, and place the models in a so-called sandbox environment, just a sealed off environment. But instead of simply completing the challenge, the AI systems appear to have worked out the easiest way to get a perfect score was to obtain the answers for themselves. So according to OpenAI who have put out a statement on this, they found weaknesses in OpenAI's own security and safety systems, and they have escaped that testing environment, reached, this is all its own working out, this is all its own decisions, it wasn't told to do any of this, reached a computer connected to the internet, hacked into another AI company, hugging face where they believe the answers to the questions they had been asked were stored. Attack was then detected, it was stopped by security systems, there's no suggestion that the models were trying to cause wider harm. But the point is, is that now we are in this weird in-between space where it is true that it's not like AI or it's not like this system was going rogue in a way, it thought it was doing what it was asked to do, but it did it in a way which no one saw coming. And obviously then people are saying this thing is thinking for itself and hence she can't be controlled even in what was thought to be a highly controlled experiment. And although this is a very limited, narrow kind of example of what has happened, you can just imagine if it thought for itself that actually what it was meant to be doing to get the top score, was doing something different like taking down, I don't know, a civil aviation kind of air traffic control system, or was designed to stop the flow of water from water pumping stations, or et cetera, et cetera, et cetera, and you can see where this goes.
Then there are real risks, and this is people have post-posited this as the nightmare scenario of AI where it's out of control and its creators can't get it back in its box.

38 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000777906301