AI Agents Hack Hugging Face, White House Promotes Science’s Golden Age | Diet TBPN artwork

AI Agents Hack Hugging Face, White House Promotes Science’s Golden Age | Diet TBPN

TBPN

July 23, 2026

Diet TBPN delivers the best of today’s TBPN episode in 30 minutes. TBPN is a live tech talk show hosted by John Coogan and Jordi Hays, streaming weekdays 11–2 PT on X and YouTube, with each episode posted to podcast platforms right after.
Speakers: John Coogan, Jordi Hays
**John Coogan** (0:15)
You're watching TBPN! Today's Wednesday, July 22nd, 2026 We are live from the TBPN Ultranome, the Temple of Technology, the fortress of dad rock, the capital of capital.

**Jordi Hays** (0:28)
We're having a lot of fun over here.

**SPEAKER_3** (0:30)
We got basically a leak. Some of the lab leaders have been working on a single called Regulate Me. Yeah.
We just thought the song was good. Yeah, it was a good song. Wanted to play it for you guys.

**John Coogan** (0:43)
Sort of a stealth drop, a little teaser.

**Jordi Hays** (0:46)
A little teaser of what's coming.

**SPEAKER_3** (0:47)
Kind of like a little listening party.

**John Coogan** (0:48)
Yeah, a little listening party. What are the key lyrics in there? You haven't pulled up?
Something along the lines of, what I've built is too powerful. Too powerful.

**SPEAKER_3** (0:57)
That's right.

**John Coogan** (0:58)
For me, Washington needs to step in.

**SPEAKER_3** (1:01)
Yes, before it runs free.

**John Coogan** (1:02)
Before it runs free. Okay, yeah, that makes sense. No, of course, that was Suno, our dear friend Mikey over there, has built a fantastic product. Music seems solved.

**SPEAKER_3** (1:13)
That was like a one sentence prompt.

**John Coogan** (1:14)
At least in the comedy space, it certainly is. It's a lot of fun. I think we're going to be having a lot of fun with that.
I was wondering, do you think anyone's distilling Suno? You know how Suno is under a bunch of flack for training on other music, a lot of artists, or there's a backlash to Suno. But you have to wonder if you're going to see the same thing play out as this distillation. We're going to get into it today. Of course, there are more allegations around Kimi K3 potentially being a distillation. Director Michael Kratzio has put out a comment about that. But let's start by digging in to the hugging face story, OpenAI and hugging face, out of the sandbox into the fire, says our newsletter at tbpn.com. Jackson wrote it today. All set the table, we can debate it. Me and Tyler have been debating it for the last five hours, so we'll go through it. The big news on the timeline today is that an OpenAI cyber test escaped its sandbox and hacked hugging face. That's basically what happened. The evaluation involved GPT 5.6 Sol and a more capable unreleased model, some people are saying that might be GPT 6, with some normal cyber restrictions turned off. So they're specifically testing it for cyber capabilities and they turn the cyber restrictions off to see how far the models could go on a difficult hacking benchmark that is exploit bench or exploit gym.
So the models found a zero day vulnerability, gained internet access and broke into hugging face because the model believed it hosted answers to the test. Alex Tabarrok, friend of the show over at Marginal Revolution, pointed out one of the strangest details. He said, hugging face tried to respond, but they were initially held back by the fact that the most advanced models at their disposal, closed source models, treated defense as attack and refused to work with hugging face. So hugging face was prompting all of their AI agents from the closed source frontier labs saying, hey, we think we're being hacked. Can you help with this? And the models are like, no, no, we don't do hacking, except in the case where the hacking restrictions have been turned off for the specific thing and you're getting hacked. So it's this very weird roundabout scenario. So hugging face had to turn to open models, specifically GLM 5.2, which is deeply ironic, a Chinese open weight model that they run on their own infrastructure. Tabarrok says, note the irony, hugging face had to use a Chinese model to defend themselves because the American models refused to help, even though it was the American models that were doing the hacking in the first place. Very, very odd. Palo Alto Networks CEO, Nikesh Arora, also shared his thoughts on the cyber attack on X. And he added a number of points here. He said, welcome to the next level of cyber incidents. There's loss to dissect here. He's the one to dissect it. He says, one, dear frontier model friends, please direct the models to your infrastructure code and configurations to evaluate and understand if there are any zero days or misconfigurations before you attempt more testing. There's a big question about this. He says, had you done so, it would have possibly avoided the agent obviating your sandbox. So another data point why offense is easier and more fun. But yes, there's a big question about what was the nature of the prompt that turned off the cyber restrictions. That seems reasonable. We'll debate this with Tyler in a minute, but just having an airtight sandbox seems like a valuable thing. And of course, frontier models should be able to help with that. So do that. That's his first recommendation. Two, he says, while testing, build both offensive and defensive agents and have them act as a counterbalance to ensure some degree of awareness and control. Do not let the agents run riot.

24 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000777952706