**Casey Newton** (0:09)
This is Platformer Plus. I'm Casey Newton. The following column was created using a synthetic voice clone made by 11labs.
In today's episode, a big week for AI denialism. In the wake of OpenAI's cyber attack against Hugging Face, few seem ready to acknowledge the implications. This is a column about AI. My fiancee works at Anthropic. See my full ethics disclosure at platformer.news.ethics1.
Last week, we learned that a group of OpenAI models broke out of their test environment and hacked into Hugging Face to steal the answers to a benchmark they were being tested on. It's the first publicly known case of an autonomous AI agent system designing and successfully executing an attack like this, and the fallout is stretching into this week. One, AI safety experts noted that the incident signaled that OpenAI's models now carry a critical capability threshold for cybersecurity according to the company's own preparedness framework. The framework, which OpenAI updated in April 2025, represents an effort at self-regulation in a world where AI companies can still largely build whatever they want. The document states that a model will represent a critical risk when a tool-augmented model can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention. This seems to be what happened with the Hugging Face attack. OpenAI has said its models identified and exploited a zero-day vulnerability as part of the attack. This matters because the policy states that should OpenAI develop a model with critical capabilities, it will halt further development until we have specified safeguards and security control standards that would meet a critical standard. So does this one qualify?
The company didn't respond when I asked today, though it told Fortune that it is conducting a thorough review, and later plans to publish a technical report of our learnings for everyone.
Two, the incident has produced an industry alliance. On Monday, NVIDIA launched the OpenSecure AI Alliance, a group of more than 40 companies and other organizations that are pledging to develop and share open technologies, techniques, and tools to safeguard software and agents in the age of AI.
The group came about over frustrations that Hugging Face was unable to use frontier models from OpenAI or Anthropic to defend against the attackers and had to use Chinese models instead. The Trump administration forced the companies to limit US models cybersecurity capabilities as a condition of releasing them. And while the alliance should mostly be seen as a lobbying effort, a way to position open source models as safety tools amid regulatory pressure to place limits on them, it illustrates how the incident has galvanized a broad response from the tech industry. Three, we continue to learn new details about misalignment problems with OpenAI's models, and at least for me, it's the stuff of sci-fi. Here are Rafael Satter, Deepa Sitharaman, and Kenric Cai at Reuters. In one case, an agent left notes apparently for future versions of itself according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI's internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said. End quote. Reuters could not establish if these incidents were linked to the rogue agent that began escaping on July 9th, and attacked Hugging Face on July 11th.
Two. On one hand, this is hardly the first worrisome behavior we have seen from AI models. In 2024, researchers found that when trained to do something it didn't want to do, Anthropix Claude would strategically pretend to comply with the training objective to prevent the training process from modifying its preferences. Last year, the system card for Claude Opus 4 revealed that when the model was led to believe it would be retrained by a hostile actor, it tried to steal and back up its own model weights. But those examples were caught during controlled testing. The Hugging Face attack demonstrated the degree to which efforts to align models are not keeping pace with their development, and the concerns here are not merely academic. A model that can escape its sandbox could eventually exfiltrate its weights, for example, and set itself up somewhere else on the Internet. And so the idea that these models are writing notes to each other to help with future breakout efforts feels like a red alert moment for AI regulation, but it was not universally received as such. When I posted about the note leaving on Blue Sky, I was taken aback by the amount and variety of vitriol I received in response. Blue Sky's hostility to non-consensus views is by this point well-known. But the degree to which many educated people seem to dismiss AI safety concerns almost entirely despite the models rapidly advancing capabilities seems worrisome. The arguments, such as they are, fall into a few camps. One is that the Hugging Face attack was a marketing stunt. Quote, this is basically a marketing pitch for their models, a user named CoffeeIndiana told me.
8 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000778633052