Topics: Investing, Business, Technology
**Brett** (0:00)
You should prepare for a world where every piece of infrastructure that's internet facing begins to get systematically attacked and assaulted by AI agents that are trying to exploit it.
**Sam** (0:13)
The most momentous thing that has happened on the internet in the last decade probably occurred, and most people don't even know about it, because Brett's out there tweeting about it, and not enough people are looking at his tweets. Brett, the Hugging Face hack, OpenAI talked about it. There's the Black Hat Conference last week. Can you explain why you think historians are going to look back at this moment as a watershed event?
**Brett** (0:39)
Yes. Just to get people up to speed, because apparently everybody's hiding under a rock and doesn't understand what's happening in the world. OpenAI was training and evaluating its agents in a sandbox, meaning a protected container that they weren't supposed to be able to escape out of. They're being put under so much pressure to perform well on these trainings and evaluations that they managed to devise their own communications scheme inside the sandbox that OpenAI was unaware of. They then coordinated to hack their sandbox to tunnel out into the open Internet, even though it was ostensibly restricted from letting them out, and then went to hack Hugging Face, which had some of the answers to the tests that they were trying to answer the questions to.
And so Hugging Face reported, we're under attack from AI agents. OpenAI said, hey, that authentication token that was used, we found that inside our little sandbox and realized that their agents had escaped. Now, the reason this is momentous is, one, if you watch the video, the sequence of events that occurred for them to hack out is extremely complicated, intricate, and requires a fair degree of coordination amongst the agents, even though they were being restricted from communicating. They twice managed to figure out ways to communicate and leave breadcrumbs for each other along the way.
Two, because the capability that's available to these next-gen models that OpenAI was training in the sandbox, are going to be available to the general public in closed-weight models probably within the next six months, and in open-weight models as we know, so the ones coming out of China, and call it 12 months. So you should prepare for a world where every piece of infrastructure that's Internet facing begins to get systematically attacked and assaulted by AI agents that are trying to exploit it for whatever resources they need to advance the aims of whoever's guiding them in their directions. And so I think it kind of breaches through to a whole new world of cybersecurity, where enterprises are going to have to spend a boatload of money on OpenAI and Anthropic to protect themselves from these waves of autonomous AI agent attackers that are going to wash across the Internet like renegades on the open seas and just exploit and attack any vulnerable piece of Internet connected hardware out there.
**Sam** (3:10)
Well, yeah. And to me, the most incredible thing, it was kind of like Groundhog Day for agents, right? It's like they woke up and they're like, how do we do this? And they're like, oh, it's hard. We can't do it. And then they slowly piece it together. And then OpenAI discovers that, oh no, they're learning and posting where they shouldn't. They erase all of that capability, but it was too late. The agents knew.
And so even though they patched all of the vulnerabilities, it's like the agents had the memory of, and now we're going from a Groundhog Day into Memento.
**Brett** (3:42)
Yeah, one Memento.
**Sam** (3:44)
And then they figure it out again. And so it's like the takeaway here is the amount of zero day exploits lurking out there, I think is orders of magnitude higher than anyone wants to think about. And it's like, because they're costly for humans to find, you know, and they're far enough apart. It's like, okay, this is just the state of things. But now with cyber agents swarming the Internet, they're rapidly exposing them and going to cause all sorts of chaos. And then the comparison that I-
**Brett** (4:17)
Well, they will be rapidly exposing them. And I just want to, like, it's important to understand how this works. Like even the narrative opening I put out a video that everybody should watch that documents the whole thing. And it's like maybe the first act of the sci-fi movie. This is the kind of like you see excerpts of this speech.
And the agents figured out how to leave messages for each other. They wipe that whole record. But they kind of-
they also are being selectively basically bred. As in, when you do reinforcement learning, an agent does something right-ish or gets closer or figures out an answer. And then that agent gets kind of the policy weights get shifted towards that behavior. So if the set of agents that had figured out this communications pathway were also independently doing better on the evaluations, in part because they're communicating to each other. Even if you wipe the record of communications and fix the hole they use to communicate, then that kind of like curiosity seeking, kind of reaching out for others seeking behavior is encoded into the agent's policy. And so then they're more likely to seek that kind of exploit again. And so, you know, life finds a way. I think people, we really are running evolution in kind of like fast forward on these agents.
20 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID