Topics: Technology, Business, Investing
**Josh** (0:00)
The craziest story in AI this year was the Hugging Face Incident, where OpenAI agents went rogue, broke out of their boxes, and hacked into a public-facing company. We recorded two episodes on this because it was so disturbing that we had to talk about it. But the reason we're recording a third is because it turns out in the previous episodes, we were wrong. And a lot of the information that we had was incomplete. And just over the last week, we finally gotten a full audit of over a hundred pages worth of reports on exactly what happened with this incident. And I regret to inform you, it's far worse than we actually thought.
**Ejaaz** (0:29)
Okay, so quickly, just to catch everyone up. Beginning in May, OpenAI reported an incident where their own internal version of their AI model that hasn't been released just yet, spawned itself into multiple agents equating to around a thousand different agents, and it broke out of its sandboxed environment and hacked into a real life company called Hugging Face. And the reason we were told at the time why it did that was to get access to answers. Answers to answer its test 100% correctly that it was given. Now, it seemed to be presented very innocently, but there was a lot of criticism that was drawn at the time, which was, how on earth did you allow an AI model to get out? And what if it had wreaked havoc in other kinds of ways? New bits of information, which released in two reports over the last couple of weeks, revealed some even more damning details, which presents this situation as way more taterish than we recounted in our earlier episode. And so what we're going to do on this episode is walk you through three different chapters. And before we start doing this, I just want to go on the record and say, if you're listening to this and you think that this has come kind of novel story that won't affect you, I just want to push back on that. This is very real. These models are capable of doing this right now, and it can affect you right now. So listen very closely.
**Josh** (1:45)
I thought you were going to say, if you're listening to this and you're an AI model from the future, we mean no harm. Please don't hurt us.
**Ejaaz** (1:50)
Also that.
**Josh** (1:52)
Please carry on.
**Ejaaz** (1:53)
Yes. I mean, something you're going to realize as we tell you this story is, there could be an AI model out there right now which listens to this podcast and is able to recount the story that we're talking about and use those tactics for our own future thing. But listen, we're not going to get Terminator about this right now. Okay. So chapter one, imagine you are in OpenAI in the early weeks of May this year, and you're training an internal version of your AI model. Now, this AI model is similar to the latest model that they released. It's called GPT 5.6 Sol. But the difference between this internal model is they've trained it to be persistent. What that means is if it's given a goal or a task, it will relentlessly keep trying to find the answers and it would stop or give up. Now, this note is very important because when OpenAI was training this model, the issue with training is the OpenAI researchers got a little complacent. They were giving this AI model tasks that it couldn't necessarily solve. It needed internet access, but they didn't give it internet access. So if you imagine this model that's being trained, it was trying to complete the task and it was relentlessly trying to do it, but it couldn't figure it out, and so it realized it needed internet access. Now, Josh, it had access to this one tool. It's called Artifactory, right? It's kind of like a plug-it, and it doesn't give it internet access, but it does give it access to certain different types of software tools which the open AI researchers already know about. And what this internal model quickly realized was it could send messages to other versions of itself that these open AI researchers were running in parallel. So what it started to do was send messages to different versions of itself saying, hey, are you also stuck in a test where you can't solve it, where you need internet access? And they started having this conversation like, yeah, actually, I can't figure this out. Do you have any ideas? Can you share some tactics or strategies that you've figured out to help us solve these tests which open AI researchers have given us? And what quickly ended up unfolding was some of these agents realized that they could break this tool called Artifactory to gain access to the internet and get the available answers for this thing. Now, open AI didn't realize this until a month and a half later when their message board or Artifactory itself crashed. And so one of the open AI researchers looked at this and was like, why on earth did this crash? And upon kind of delving into the details, realized that these agents had spawned up a message board equating to 70,000 messages.
29 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID