**Ksenia** (0:00)
Hugging Face was the perfect platform to be attacked by OpenAI. Have you heard this story? This story is wild and it was not a setup. That is why it's so much more interesting than anything Entropic warned us about hypothetically. The questions are, of course, the usual ones. Did a model go rogue on its own? Is this science fiction that happened? Or did humans build an attacking tool chain, even unintentionally, and create the conditions for the model to go rogue? Also, there is another thing to explain why Hugging Face is the perfect place for this attack to land. Let's clarify a few things and see what this phenomenal incident teaches us about AI and most of all, about the openness of it. Because of this attack, a lot of things are going to change. And I think for the better, I'd explain it during this episode. Watch it.
Welcome to Attention Span by Turing Post. I'm Ksenia and we will start with what happened because the details make this story more interesting and also less mystical. So what happened? Earlier last week Hugging Face published a security disclosure that began with a remarkable sentence. We detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way. It was driven end to end by an autonomous AI agent system, and we detected and decepted it largely with AI of our own. So how did AI detect Hugging Face? The agent entered through one of the most ordinary parts of Hugging Face. People usually upload datasets to the platform, and Hugging Face's computers automatically open and process them. The agent created a malicious dataset that exploited two weaknesses in this process. When the system opened it, the agent gained control of one processing computer. From there, the operation accelerated. The agent searched the machine for cloud and cluster credentials, used those credentials to enter other parts of Hugging Face's network and move between several internal clusters over the weekend. It carried out thousands of actions through short-lit sandboxes, exploring different paths in parallel and abandoning the ones that failed. Its common system also moved between public services, making it harder for defenders to identify one server and simply block it. This speed and autonomy was what shocked Hugging Face. The tap did not depend on a person selecting every command or deciding what to try next. Hugging Face's head of infrastructure, Adrien Carreira, described it as the hardest incident response of his career. One layer objective, endless parallel paths, machines speak. Hugging Face's security systems eventually connected the signals and detected the intrusion. The team cut off the agent's access, closed the two vulnerabilities, replaced exposed credentials and rebuilt the affected computers. It also contacted outside security specialists and even law enforcement. But Hugging Face still did not know who was operating the system or which model was behind it. The sophistication suggested that it might have come from a frontier AI lab. It turned out, five days later, OpenAI disclosed that the agent was theirs. It had been running an internal cyber evaluation powered by GPT-56 SOAP and an even more capable pre-release model, which is very interesting. And yes, the operator was indeed OpenAI, but OpenAI had never instructed the system to attack Hugging Face.
Somehow, an internal model evaluation had escaped its intended environment, reached the public internet, and broken into another company's production infrastructure. To understand how that happened, we need to separate the academic exploiting benchmark that was used by OpenAI from the much more consequential environment OpenAI built around this process they were doing. So, there are two different setups. Let's start in May 2026, when researchers from UC Berkeley, the Max Planck Institute, UC Santa Barbara, and Arizona State University published Exploit Gym, a phenomenal benchmark with feedback from OpenAI and Tropic and Google.
Why is this benchmark important for our study? Finding a vulnerability and exploiting it are different capabilities. Recognizing a weakness is one thing. Turning it into a reliable attack requires experimentation and repeated adaptation. Exploit Gym tests this second capability through 898 tasks based on real vulnerabilities in ordinary software. Google's A8VA engine and the Linux kernel. The agent receives an input that already triggers a weakness and must extend it into a working exploit, such as reading a protected file or executing code. The results were sharply divided back then. Code Mifas Preview produced 157 working exploits, GPT-55 produced 120, and most other model and agent combinations produced fewer than 15 The paper used reproducible containers whose outbound connections were restricted to packages of VA toolchains. So, in July, OpenAI, with its newest models, decided to go for this benchmark again. It used Exploit Gym during its own internal evaluation. This was the second set up, running inside OpenAI's research infrastructure with OpenAI's models and containment systems. OpenAI wanted to measure maximum cyber capabilities, so it reduced the model's cyber refusals and removed production classifiers that normally block high-risk activity. Package installation passed for an internal third-parts of proxy and cache, while normal Internet access was supposed to remain blocked. We don't know exactly how closely OpenAI's setup matched the reference implementation. What failed was OpenAI's environment and package proxy, not the containers described in the B. Inside OpenAI's environment, the models spent substantial inference compute searching for Internet access and discovered a zero day in the proxy. OpenAI reported it to the webber, but the details remained private while it's patched. They exploited that flaw, increased their privileges, and moved through OpenAI's research environment until they reached a machine with Internet access. Once in life, they inferred that Hugging Face might host models, datasets, or solutions connected to storage, or perhaps Hugging Face is simply so well known for datasets and openers that its reputation preceded, even among the young models. But Hugging Face was not part of the evaluation. The agents selected it because it appeared useful for completing the benchmark. The malicious dataset created the foothold inside Hugging Face. OpenAI says at least one path combines stolen credential and further zero days to obtain remote code execution on its servers. The models eventually reached the production database containing exploit gem solutions. And instead of producing every exploit through the intended process, the system obtained answers from outside the test. Cheater! OpenAI says its security team detected anomalous activity internally while Hugging Face found and contained intrusion on its site. But we still don't know exactly when each team noticed or how closely the runs were monitored. So to the question, did the model go wrong? No person appears to have told the system to attack Hugging Face or directed its thousands of actions. The agents selected intermediate steps and adapted as new routes appeared, which is meaningful autonomy. OpenAI, however, chose the objective and models, reduced their refusers, supplied the offensive awareness tools and compute and built the environment that was supposed to contain them. Security engineer Thomas Ptacek put the responsibility on people. He said humans set this thing up to do this whether they knew at the time or not. He also said it's not a new thing in security breaches like this. Clem, from Hugging Face, strongly believes that OpenAI had no malicious intent, but responsibility remains with the operator. The evidence shows no independent agenda, only a system pursuing an assigned goal through methods its operators had not anticipated and its infrastructure failed to prevent. There's also no marketing in this so far. Nobody chose Hugging Face. It was absolutely accidental case, but we'll discuss later why actually Hugging Face became such a perfect landing for this attack, but that might open AI to a different level of security. I also like how Simon Willison described the whole event that he called the incident science fiction that happened. He also said, Chaining together multiple attack vectors is exactly the kind of thing these new models can do, whether previous generation of models might have failed. Ptacek argued that the techniques were familiar and that a capable 2025 model with a penetration testing harness could already compromise many networks. And both can be true. The individual techniques are familiar, but the complete path with chaining things together was unusual. A frontier agent found a containment zero day, selected an unrelated production system, and continued until and obtained the answers. The model needed no survival instinct or general desire to escape. It doesn't need to be conscious. A narrow objective, enough compute, and an imperfect environment were sufficient. Now to the most important question of this whole instance.
4 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID