Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan artwork

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast

June 22, 2026

AI Engineer World’s Fair regular bird tix will sell out ~today! Join us next week ahead of the Late Bird price hike and get >$40,000 in sponsor credits for attending!
Speakers: swyx, Zico Kolter, Matt Fredrikson
**swyx** (0:04)
Okay, we're here in the studio with Gray Swan, Matt and Zico, welcome.

**Zico Kolter** (0:08)
Great to be here.

**Matt Fredrikson** (0:09)
Yeah, thanks for having us.

**swyx** (0:09)
You're visiting from Pittsburgh?

**Zico Kolter** (0:11)
That's right.

**swyx** (0:12)
The home of all good computer science. I don't know if I'm overstating things.
Very strong university.

**Zico Kolter** (0:17)
Yeah, CMU has been the center of a lot of AI since really the dawn of the field.

**swyx** (0:22)
Yeah, especially a lot of self-driving, some language learning. Congrats on your Series A. You're here because you're attending Snowflake Summit and Snowflake is one of your investors. Let's introduce Chris Lee at the top.
What is Gray Swan and what have you chosen to be your startup domain?

**Matt Fredrikson** (0:42)
Yeah.
At Gray Swan, our mission is to empower everyone to use AI safely and securely.
Really, large language models are, at the end of the day, software. If you want to deploy them, build applications on top of them, you need to be aware of what the vulnerabilities might be, what can go wrong, and not just in everyday use, like you're innocently using an agent and maybe it makes a mistake in a tool call, but also in worst-case scenarios where there might be an attacker who has an incentive to make your agent misbehave, leak data, steal credentials, things like that. Gray Swan really grew out of our research. Zico and I have been at Carnegie Mellon for some period of time, over a decade, looking into just this. What are the new vulnerabilities and attack surfaces in especially deep learning systems? How do you test for them? How do you understand the scope of how severe they can be? Once you know that there is a vulnerability, there is a problem, how do you fix it? How can you do inference more robustly? What can you put in place to make sure that these bad outcomes don't come to pass?

**swyx** (2:05)
Honestly, a very fruitful area of study for any academic. Throwback, this is 10 years ago, which is literally the entire time. I actually got a lot of inspiration from Ian Goodfellow, who's a friend of the pod.
This is one of those initial adversarial settings.

**Matt Fredrikson** (2:24)
This paper was directly inspired by Ian's. Yeah.

**swyx** (2:29)
Zico, what about your side of the story?

**Zico Kolter** (2:31)
Yeah. So like Matt, been faculty at Carnegie Mellon for a while, I think fundamentally, look, I think that in some sense we're all here because we believe in the transformative power of AI and we think that this has already transformed the way the entire software ecosystem works and it will transform how many other ecosystems work going forward.
The issue though is that these systems just fundamentally behave very differently from software we're used to. I don't mean in terms of AI can find vulnerable software though it can also do that and is also transforming that. I just mean that AI systems have inherent different types of vulnerabilities. They can be tricked like people get tricked sometimes. And so you need a different mindset about security when you're thinking about AI systems. And especially when there's the possibility of correlated failures. So it's not just that there's a lot of AI systems out there. It's that there's actually a few models that everyone is using. And if you find vulnerabilities in the agents that everyone uses, things like Codex and Claude Code, you can actually start to now essentially have a new exploit, a new class of exploit.
Fundamentally, I think there has to be a different mindset about the nature of AI security as there is for traditional security.
And while a lot of that's going to, of course, happen at the AI companies themselves, labs themselves, there's also a real value. And of course, to be very clear, the labs are doing a lot of work in these areas. But there's, just like in most domains, when a new platform emerges, it's very common for there to also emerge a security system separate from it, right? In addition to it, as a separate service that's provided. And I think that's where we are right now with AI. And I think there's a need for specifically minded AI safety and security providers. There's a demand for this, and there's going to be much more demand for this coming up. And that's why it felt like a really good time to sort of focus on this problem, both in research, because we still do research on this topic too. And we're continuing research actually at Gray Swan, but also in terms of a commercial offering.

**swyx** (4:55)

59 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000773793683