Topics: Technology, Business, Entrepreneurship
**Joel de la Garza** (0:00)
One of the interesting things in the OpenAI hugging face breach has been the difficulty that hugging face actually had responding to the incident.
**Nick Warner** (0:09)
Model providers have great reason to establish guardrails, safeguards because these are super capable systems. The unfortunate side effect of that is, as a defender, I may not be able to respond effectively.
**Max Pollard** (0:22)
The challenge with the existing security tools that are out there is they really were built to tackle two things. The first being people and the second is malware. AI and AI agents and agentic processes are neither one of those things.
**Nick Warner** (0:32)
Even some of the more modern techniques like deception that work really well.
**Max Pollard** (0:36)
It's ironic. We're defending AI, we're also defending from AI. Fifty percent of enterprise apps will be agentic by the end of this year, and the average enterprise is something like six or 7,000 unique pieces of software within their environment. The problem is going to get more complex and more challenging.
**Joel de la Garza** (0:51)
We seem to be speed running every technology cycle that's ever happened before this one. What is the path forward for a lot of this environment?
**SPEAKER_4** (0:59)
Cybersecurity was built to defend against two things, people and malware.
AI agents are neither.
In this episode, a16z's Joel De La Garza sits down with Nick Warner of Neo and Max Pollard of Cotool to unpack what that means for security teams as increasingly capable models move from the cloud, onto endpoints, and into enterprise software.
They discuss why AI guardrails can actually make life harder for defenders, why traditional signatures and behavioral detection are starting to break down, and what happens when software no longer behaves predictably enough for security teams to define what normal looks like. And from Black Hat, they look at the other side of the equation. The same AI that's creating an entirely new attack service is also giving defenders tools they could never have built before.
**Joel de la Garza** (1:52)
Awesome. Well, thank you so much, guys, for joining us. I think maybe let's set the stage for the discussion. It's been a very active couple weeks.
We've obviously had the crazy pace of AI development. Every week, there seems to be a new model release, there seems to be a new capability, there's some new fields metal getting one, or some new vulnerability getting discovered. And in the news recently has been the report that models from Frontier Labs have found a way to escape containment and hack things on the Internet, which has been a pretty interesting revelation. That's a very sophisticated capability. I guess there's probably two things happening, right? There's a discussion about, well, how secure is the Internet actually? Also combined with, man, these models are surely progressing and doing some great stuff. So we've got the two of you here to discuss this, and I think maybe Max will start with you.
One of the interesting things in the OpenAI hugging face story breach event has been the difficulty that hugging face actually had responding to the incident. And so maybe you could tell us a little bit about sort of like what was happening there. Why won't these models help the good guys? What's going on and what's the deal?
**Nick Warner** (2:57)
Yeah, so, I mean, super topical, like the model providers have great reason to establish guardrails, safeguards, because these are super capable systems and we don't want attackers using them for nefarious purposes. And so they have a responsibility to make sure those guardrails are in place. The unfortunate side effect of that is as a defender, let's say you're triaging an incoming bug bounty report or something of that nature, often times you're going to be asking very similar questions to a probing attacker, which is, hey, what piece of this software is vulnerable? How would I exploit it? Can you validate this, right? And so part of the side effect of these model providers doing their job is, as a defender, I may not be able to respond effectively. Now, for Hugging Face specifically, they had the luxury of having open-laid and open-source kind of in their DNA. And so they were able to, when they saw these cyber refusals or these guardrails being triggered, fall back to GLM 5.2 in their case, but could have been Kimmy K3 or a Quinn model, something that isn't going to have those guardrails in place.
And so for defensive teams, flexibility is kind of becoming paramount, right? You need the ability to kind of fall back in the case of refusals.
**Joel de la Garza** (4:11)
Yeah, absolutely. And I guess the question would be, what is it specifically refusing? Like, I guess the thing maybe for folks that are listening is, like, I get that it stops you from saying, hey, go hack this website. Sure. I mean, go to citibank.com and change my account numbers, right? And like, the refusal makes sense there. But what are Blue Team's doing that is specifically generating the refusal? Why does it look like a hacker to the models?
18 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID