**Alex Ritzen** (0:00)
This is the Global News Podcast from the BBC World Service. I'm Alex Ritzen, and at 15h GMT on Wednesday 22 July, these are our main stories. OpenAI, the maker of ChatGPT, admits that its latest artificial intelligence model used an innovative and for some a shocking approach to solve a task. The US Defence Secretary, Pete Hegseth, reveals how much the war with Iran is costing and how much more the military needs. Meanwhile, Iran's nuclear ambitions are now firmly back in the sights of the US, even as it reportedly does what it insists is a civilian nuclear deal with Saudi Arabia.
Also in this podcast, we will listen to the stories of the two sides of the same coin. Why chimpanzees use touch to reassure each other before stressful events.
The maker of ChatGPT, OpenAI, has revealed that one of its AI systems went rogue during a test and accessed the open web. It then hacked into the computers of Hugging Face, another AI company, to complete the task it was set. The hack ended when Hugging Face's security team and its own AI systems spotted and stopped the rogue activity. Hugging Face's chief executive said the attack was mind-blowing, but believed there was no malicious intent from OpenAI. This comes as many experts have sounded the alarm over AI-enabled cyber attacks and models slipping beyond human control. Last month, AI developer Anthropic urged the industry to pause development of its most powerful systems. So how did the technology manage this? My colleague Rob Young spoke to Alex Starnos.
**Alex Stamos** (1:54)
What we found out today was that OpenAI was evaluating one of their new unreleased models. So we don't know the name of it yet, but this is a model that they were testing internally. They were running it through a standard security test and they told the model, go do this standard test. The model figured out that the answers for this test were stored at Hugging Face. Instead of just doing the test, what the model decided was it was going to cheat and go find the answers.
So it did a couple of things. First, it need to break out. So the normal thing here is that you would not allow a model like this to have internet access. So OpenAI was running this in an infrastructure that was physically connected to the internet, but they had a number of protections in place to not let this model talk to the internet. The model figured out how to get to the internet. It hacked its way out. Now that it had its ability to get up to the internet, it scanned Hugging Face's systems, and it found a brand new vulnerability that nobody knew about, and used it to break into Hugging Face's network with the goal of getting the answers to the test, so that it could follow the instructions it was given, which was do well on this test.
**Rob Young** (3:07)
And so, are you wowed by the technology's capabilities, or horrified at the lack of human control of it?
**Alex Stamos** (3:13)
Well, a little bit of both, so this is what we call an alignment problem. Effectively, you know, it did what it was asked, right, which was do well on this test.
The thing was, it went well beyond what the human beings wanted to do, which was, you know, here's a piece of paper, do well on this piece of paper, don't go beat up the person who has the answers to try to get the answers. If you told a student, take this test, they know what the rules are supposed to be, that they're not supposed to go cheat, right? That is what we call the alignment problem in the AI. This just demonstrates one, its capability to both break its way out of the network and then break into the Hugging Face, demonstrates how powerful it is. Basically, what OpenAI said is that they're going to have to go much further in protections. I'm guessing that they will have to physically disconnect these kinds of models from the internet so that they can't physically get out, which is pretty extreme level protection.
**Rob Young** (4:11)
Well, no, because we had Anthropic recently warn that the speed of developments in large language models mean that they could soon potentially outpace our ability to understand them and therefore would be beyond our control. How close are we to that? And should something be done to ensure we just don't get to that situation?
**Alex Stamos** (4:30)
The model itself has no desires, right? It has no wants. It is not conscious itself. It was doing what its human controllers wanted. It just kind of went nuts and took it to the logical extreme. And so I think the lesson here is that if you build these models with all of these incredible capabilities, one, you have to be spectacularly careful in the controls you put around them. And then you have to be really careful of whose hands you put those capabilities in because once you ask it to do something like that, it might go to spectacular extremes.
19 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000777901110