Topics: Politics, News, Technology
**Jordan Schneider** (0:01)
What does the Hugging Face hack portend? Are we in for a cyber apocalypse? To discuss, we have on Truly the Perfect Guest, Joshua Saxe spent his formative years mucking around in the deep state as a contractor for Darpa and the NSA before spending the past four years at Meta, ending as its AI Security Tech League.
He recently branched out on his own as the CTO and co-founder of New Startup. For starters, what was the Hugging Face hack?
**Joshua Saxe** (0:30)
OpenAI was training a new model that is capable of solving long-horizon tasks, so the tasks that take many hours to complete.
The way that kind of training works is you train thousands of copies of a model in parallel to solve tasks like that. You have the model to go out and try to solve a set of tasks. You see which one succeeded and which one's failed, and then you update the model's parameters. You sort of update the models so that they do more of the good behaviors that led them to solve the task and less the bad behaviors. You just keep doing that iteratively at a very large scale over the course of usually weeks or sometimes months. They were doing this kind of training.
While the models are being trained, they broke out of their training environment, hacked internal infrastructure at OpenAI, and then they eventually hacked another company, Hugging Face, which is an open-source model and model evaluation repository. It's widely used in the ML community. The models hacked into Hugging Face, and at a certain point Hugging Face discovered what had happened, and eventually everybody figured out what had happened, that these models had broken out. This whole thing unfolded over the course of many weeks, actually. I think there were three different cases. So there was a couple of early cases where the models broke out of their containment environment just at OpenAI, and that was detected in OpenAI. I sort of thought they fixed the problem. It turned out they didn't, and then in the third case, the models actually hacked Hugging Face.
**Jordan Schneider** (2:15)
The world is currently freaking out right now. Are they correct, too?
**Joshua Saxe** (2:20)
I think for those of us who've been around AI safety for a long time, it wasn't that surprising. Like, I didn't find that surprising. But I think it is alarming. It's alarming that it happens, because we could have predicted this, and not just OpenAI, but also Anthropic and Meta, and also a testing company called Irregular, and also the UK AI Safety Institute all had incidents like this happen in the last year or so.
As an industry, we knew this could happen, and the basic measures that should have been taken to prevent it weren't taken. So that's troubling. We need to fix that.
**Jordan Schneider** (3:07)
How easy is this as a thing to stop, and what is the incentive structure that's wrong in all of these organizations, including ostensibly very safety-pilled ones, for them to not have been able to make sure that their latest models didn't run amok and in customers' servers?
**Joshua Saxe** (3:28)
Yeah.
So I'll say a couple of things. So I was at Meta, and I started the team. I started the team that actually did Frontier and Cyber Capabilities evals, so I have that background. I'm just giving you a sense of where I'm coming from, because I can't speak directly to exactly what happened in the labs, because it's private. So I had that experience of starting a team that did that, and we did similar experiments to the experiments that were happening when the opening-end models broke out.
I also, I guess another point of background is we all kind of know each other in this community, or like many of us do, so across the labs. So with that said, I would say that the culture among the training teams and the EVALS teams at the labs is there's a Wild West kind of feeling to the whole thing, like where everybody's just under a ton of pressure to move really quickly. There's enormous time pressure to release new models. There's tremendous awareness of how any given lab is doing relative to the other labs in terms of how capable their models are. And there's also, the field's evolved super fast over the last four years since Chad JBT launched in fall 2022 And the whole thing has just been kind of a blur. It's been like four years, but it just feels, there's like a time compression feeling around it. Like anybody who's worked in AI in the labs in that time has been working 60-hour weeks the whole time. And just, it's not surprising to me that you could have a kind of like grad student computer science lab kind of culture around the security of these model training runs, which is I think what is the case, right? I think that's why you've seen these these these these lab escapes. You know, it is ironic because frontier AI people talk about safety all the time. But yeah, that irony is a real thing in reality. Like the security practices are not very good because of all those dynamics.
50 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID