**Sander Schulhoff** (0:00)
I found some major problems with the AI security industry. AI guardrails do not work. I'm going to say that one more time. Guardrails do not work. If someone is determined enough to trick GPT-5, they're going to deal with that guardrail, no problem. When these guardrail providers say, we catch everything, that's a complete lie.
**Lenny Rachitsky** (0:16)
I asked Alex Komoroske, who's also really big in this topic, the way he put it, the only reason there hasn't been a massive attack yet is how early the adoption is, not because it's secure.
**Sander Schulhoff** (0:25)
You can patch a bug, but you can't patch a brain. If you find some bug in your software and you go and patch it, you can be 99.99% sure that bug is solved. Try to do that in your AI system. You can be 99.99% sure that the problem is still there.
**Lenny Rachitsky** (0:39)
It makes me think about just the alignment problem. You got to keep this God in a box.
**Sander Schulhoff** (0:43)
Not only do you have a God in the box, but that God is angry and that God's malicious. That God wants to hurt you. Can we control that malicious AI and make it useful to us and make sure nothing bad happens?
**Lenny Rachitsky** (0:56)
Today, my guest is Sander Schulhoff. This is a really important and serious conversation and you'll soon see why. Sander is a leading researcher in the field of Adversarial Robustness, which is basically the art and science of getting AI systems to do things that they should not do. Like telling you how to build a bomb, changing things in your company database, or emailing bad guys all of your company's internal secrets. He runs what was the first and is now the biggest AI red teaming competition. He works with the leading AI labs on their own model defenses. He teaches the leading course on AI red teaming and AI security. And through all of this has a really unique lens into the state of the art in AI. What Sander shares in this conversation is likely to cause quite a stir. That essentially all the AI systems that we use day to day are open to being tricked to do things that they shouldn't do through prompt injection attacks and jailbreaks. And that there really isn't a solution to this problem for a number of reasons that you'll hear. And this has nothing to do with AGI. This is a problem of today. And the only reason we haven't seen massive hacks or serious damage from AI tools so far is because they haven't been given enough power yet and they aren't that widely adopted yet. But with the rise of agents who can take actions on your behalf and AI powered browsers and soon robots, the risk is going to increase very quickly. This conversation isn't meant to slow down progress on AI or to scare you. In fact, it's the opposite. The appeal here is for people to understand the risks more deeply and to think harder about how we can better mitigate these risks going forward. At the end of the conversation, Sander shares some concrete suggestions for what you can do in the meantime, but even those will only take us so far. I hope this sparks a conversation about what possible solutions might look like and who is best fit to tackle them. A huge thank you for Sander for sharing this with us. This was not an easy conversation to have and I really appreciate him being so open about what is going on. If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube. It helps tremendously. With that, I bring you Sander Schulhoff after a short word from our sponsors. This episode is brought to you by Datadog, now home to Eppo, the leading experimentation and feature flagging platform. Product managers at the world's best companies use Datadog, the same platform their engineers rely on every day, to connect product insights to product issues like bugs, UX friction, and business impact. It starts with product analytics, where PMs can watch replays, review funnels, dive into retention, and explore their growth metrics. Where other tools stop, Datadog goes even further. It helps you actually diagnose the impact of funnel drop-offs, and bugs, and UX friction. Once you know where to focus, experiments prove what works. I saw this firsthand when I was at Airbnb, where our experimentation platform was critical for analyzing what worked, and where things went wrong. And the same team that built experimentation at Airbnb built Eppo. Datadog then lets you go beyond the numbers with session replay. Watch exactly how users interact with heat maps and scroll maps to truly understand their behavior. And all of this is powered by feature flags that are tied to real-time data, so that you can roll out safely, target precisely, and learn continuously. Datadog is more than engineering metrics. It's where great product teams learn faster, fix smarter, and ship with confidence. Request a demo at datadoghq.com/lenny. That's datadoghq.com/lenny.
70 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000742197136