Reading model benchmarks like a pro, Mythos is looming, and Claude talk caveman, save big token artwork

Reading model benchmarks like a pro, Mythos is looming, and Claude talk caveman, save big token

Dev Interrupted

April 10, 2026

Is the secret to slashing your token costs by 65% forcing your LLM to speak like a caveman?
Speakers: Ben Lloyd-Pearson, Andrew Ziegler

Topics: Technology

**Ben Lloyd-Pearson** (0:05)
Andrew, what opinion have Claude Mythos?

**Andrew Ziegler** (0:10)
Andrew have many opinion, Claude Mythos. Claude Mythos, big model.

**Ben Lloyd-Pearson** (0:19)
Yeah, alright, we'll get into why we're talking like a caveman, but yeah, Andrew, Claude Mythos, it seems like everyone's talking about it. Is this what you're hearing out there right now?

**Andrew Ziegler** (0:27)
Oh, yes, indeed. And welcome to the Friday Deploy, y'all. It is true that people are talking about Mythos out in the wilds here at HumanX. I've definitely heard it on people's lips here on the expo floor, and definitely the security-minded companies have been top of mind for them. I think Mythos is a really fascinating kind of sea-change event in model capabilities, especially in a realm where there's already a huge amount of, like, disparate abilities between attackers and defenders in the cybersecurity space. Anything you're going to put out there that defenders can leverage, an attacker can leverage ten times better and faster and more aggressively, the really unbalanced world. So Mythos entering it is, you know, I think really trepidatious for some people. What have you been hearing about Mythos?

**Ben Lloyd-Pearson** (1:16)
Yeah, I mean, the cybersecurity, we'll get into that here in a moment. You know, I've just kind of come to expect that every, every, I don't know, every month now, maybe every two months, the timelines seem to be condensing where some new major incremental improvement comes out. So, you know, the current frontier models are pretty amazing. So I only kind of expect the next ones to sort of take it a level up. So, but yeah, so let's get into it. Yeah, as you mentioned, this is the Friday Deploy. I'm your host, Ben Lloyd-Pearson.

**Andrew Ziegler** (1:45)
And I'm your host, Andrew Ziegler.

**Ben Lloyd-Pearson** (1:47)
So this week, we are covering this AI cybersecurity arms race that we were just discussing. We'll also go over how to read AI scorecards and benchmarks. We'll talk about some open source frontier breakthroughs. And then we're going to get into why Andrew and I started this episode talking like cavemen, because it's actually a really cool story.

**Andrew Ziegler** (2:06)
Andrew excited.

**Ben Lloyd-Pearson** (2:07)
Yeah, exactly. But let's kick it off with Project Glasswing. So this is a new thing that has been announced from Anthropic. They're partnering with a bunch of major tech companies to use advanced AI models for finding software vulnerabilities before attackers do. So we were getting into this a little bit, but Claude Mythos has been previewed, or Anthropic released a preview of it. And along with it, it seems to be becoming a lot of just like concerns and warnings about how it could be used maliciously. One of those ways is that it's finding a lot of security vulnerabilities and very commonly used libraries. Like, for example, it found very old bugs within FFmpeg that have have been undetected, you know, despite being scanned thousands of times over over the years. And it really does, you know, highlight that there is this urgent arm race between AI-powered defenders and attackers, you know. So in this, Anthropic is partnering with like the Linux Foundation and a whole bunch of other like big logos. I don't remember the whole list, but there was a bunch of companies on that list. And they're committing 100 million in usage tokens to help organizations scan their systems. So it's really great to see them being proactive. And, you know, I think the Mythos has like some doom and gloom around, like how security is about to become a nightmare. But at the same time, I think it's also really good to focus on the good things that AI can do. Like long-standing security vulnerabilities in FFMPEG should be fixed, whether or not we have AI. So we can also use AI to do those things. So, yeah, Andrew, what are you thinking? Because I know you've been following some of this stuff pretty closely.

**Andrew Ziegler** (3:46)
The Mythos rollout, or rather development, and then them partnering with organizations in this Project Glasswing initiative, I think, is a really great play to see from Anthropic, who sees itself as a partner with the software ecosystem. The organizations that they reached out to maintain the software and operating systems that power the entire world. And by just partnering with that small, very concentrated large group of organizations, you get a really wide spread over all of the tools we use every day. So Anthropic was definitely really smart in partnering with those people first, because they acknowledge this inequality between attackers and defenders in the cybersecurity space. And I definitely think that this is... I think to the ethos, I think of Anthropic as well. I don't know if necessarily OpenAI would have made that same decision. I think they probably would have more welcomed the disruption, as opposed to trying to gently roll it out. However, of course, Anthropic 2, I think, also has ulterior motives here. We're living at a time where serving Opus to their customers is extremely difficult for them at scale right now. There's been a lot of talks about folks having different experiences with Claude Code and with Opus in particular, getting good usage out of it, getting the limits and the mileage out of it that they used to. And there have been strains around delivering that compute.

25 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID