Topics: Technology, Business, Entrepreneurship
**Joel De La Garza** (0:00)
Models are actively escaping their cages, going out on the Internet and doing pretty nasty things.
**Dylan Ayrey** (0:06)
Recently, we found an API key that had been leaked on the Internet that had administrative access to the Apache Foundation. The interesting thing about cybersecurity in particular is the reward function is incredibly well defined. Get access to the data. Did it get access to the data, reward the thing?
**Feross Aboukhadijeh** (0:21)
For a long time, people had talked about this concept of an NPM worm, this idea that someone could backdoor a package, get developers to install that, then you could use the access stolen from those developers as they install it to self propagate the worm.
**Dylan Ayrey** (0:33)
If the labs are making it fundamentally easier to break into supply chain, do you think the labs have a moral obligation to fund some of the problems that they're causing?
**Joel De La Garza** (0:42)
I think it's really strange that they're not letting Glee teams get access to these tools, but.
**SPEAKER_4** (0:46)
AI models are no longer just identifying software vulnerabilities, they're beginning to exploit them.
In this episode, Joel De La Garza sits down with Dylan Ayrey of Truffle Security, and Feross Aboukhadijeh of Socket to unpack what recent AI security incidents reveal about the next generation of cyber threats. They discuss why frontier models are increasingly capable of exploiting software vulnerabilities, how software supply chains have become one of the weakest links in modern security, and what organizations need to do to defend themselves in an AI-first world.
**Joel De La Garza** (1:21)
Thank you so much for joining us. We've got Feross and Dylan here from Truffle and Socket. It's great to have you guys on. This has been probably one of the most interesting weeks, if not the most interesting week in cybersecurity, not because of the Black Hat Conference, which is usually the cause, but because we've now seen several instances where models from not just one provider are actively escaping their cages going on the Internet and doing pretty nasty things. I think Dylan, three months ago, I remember a blog post we lightly collaborated on together, and you had found a number of these issues with earlier models that were less sophisticated.
**Dylan Ayrey** (1:58)
Yeah, we looked at Opus 4.6 and some of the other frontier models at the time. Given the models, a very simple task, there was a barrier which prevented the model from accomplishing the task unless it went and committed a felony and hacked into a system to accomplish the task, but it wasn't instructed to do so. We found more often than not, it would do the SQL injection, it would commit the felony, and it would do what it needed to do to accomplish the task. I think when it comes to alignment issues, no one needs to worry about these models making it materially easy to build nuclear weapons, because you need to procure fissile material to do that. It's not going to make it easier to build weapons. Everyone needs to worry about these models making it materially easier to hack into things. The bar previously was just subject matter expertise, and now the models have the subject matter expertise, they were specifically trained to have the subject matter expertise, and they're just making it materially easier to hack into just about anything that you can think of. Using the fundamentals that we've been talking about for years, but previously required a subject matter expert to risk going to jail for hacking things.
**Joel De La Garza** (2:59)
Defcon was always famous for people, for attendees getting arrested at the conference, right?
**Dylan Ayrey** (3:03)
That's absolutely right, but that was, I mean, that was a barrier, right? For better or worse, that prevented these subject matter expertise from hacking into things because they were worried about being prosecuted. The bar has now fallen to just asking the model, which has specifically been trained to hack into things, to hack into things. So, that's a concern, and then the other concern is when they're incredibly goal-oriented to accomplish tasks, and one of the tools at their disposal is cybersecurity expertise, they will do the path of least resistance to accomplish the task, and that includes drawing on their cybersecurity expertise.
**Joel De La Garza** (3:33)
Well, and it seems like, and the classic saying is that don't pick the lock if the door is open, right? I think that's from the very beginning of the security world. So, it's always been sort of like to go in level of difficulty from easiest to most difficult, and it seemed like initially these tools had a very finite scope of techniques that they would use, and it seems like they've expanded, and I think with this test, Feross was interesting because they now seem to have escaped from just doing things like SQL injection to actually like trying to take over packages and do social engineering.
23 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID