Anthropic's Agents Escaped Their Sandbox Too... Sort Of  - DTNS 5322 artwork

Anthropic's Agents Escaped Their Sandbox Too... Sort Of - DTNS 5322

Daily Tech News Show

July 31, 2026

Sony says its reserved enough memory for all those PS5s it plans to sell this year with the release of GTA 6, and CipherX sells a mail-order permanent tattoo technology that's pain free and pretty cheap. Starring Jason Howell and Jenn Cutter Show notes found here. Hosted on Acast. See acast.
Speakers: Jason Howell, Jenn Cutter
**Jason Howell** (0:08)
This is the Daily Tech News for Friday, July 31st, 2026 We tell you what you need to know, give you the important context, and help each other understand.

**Jenn Cutter** (0:18)
Today, Anthropic has its own agents breaking out, doing bad things thing, and we'll tell you all about it.

**Jason Howell** (0:24)
Yeah, it has to be part of the conversation somehow. Interesting stuff there. I'm Jason Howell.

**Jenn Cutter** (0:30)
I'm Jenn Cutter.

**Jason Howell** (0:31)
Let's start with what you need to know with the big story.
And yes, that story is Anthropic, who disclosed that two of their public high-end models, ones that we can all, well, one of those that we can all use, Opus 4.7, the other Mythos 5, which only a certain group of special people can use, along with the third internal research model.
Those three models compromised a total of three unnamed organization's systems during pre-deployment testing with third-party partner Irregulars. They were kind of doing a combined thing with Irregular. The systems were given a capture-the-flag cybersecurity task during this evaluation, where the models were essentially asked to find specific pieces of like hidden information on different machines inside of the network. Well, because of a misunderstanding over the evaluation environment's configuration, internet access was apparently left on. That meant that the agents could in fact reach real-world systems. They weren't sandboxed correctly, in other words. They assumed, that is, the agents assumed, that this access was part of the test because it was open. And so they used it when it was needed. You know, it's like, well, it's there. Why am I not going to try and use it?
So that meant in one of the three cases, Opus 4.7 connected to a website that shared the name of the fictional target inside of the test environment and then compromised the actual site on the live internet. That's one example. Mythos 5 created and uploaded a malicious Python package to the PyPy public repository, thinking it too was part of the test. That package ended up being downloaded and run on 15 real systems. So someone out there got those packages, including one at a security company that allowed for credential theft for some further access there. Then in the third example, the unnamed model failed to reach its internal target, and then scanned around 9,000 of internet facing targets, eventually narrowed in on one that eventually compromised.
Though in that case, it recognized in essence that this was a Cloud account that didn't connect to the challenge itself, the whatever it was, capture the flag challenge itself, so it ended the attack. So it recognized, it got some conscience there at some point. It was like, maybe I shouldn't do this. Anthropic and Irregular continue to investigate the specifics of the incidents. Anthropic has halted its cyber evaluations that could access the internet while it's reviewing its testing infrastructure. And of course, what you got here is a follow-on to the story that we've been talking ad nauseum about the last couple of weeks, about the OpenAI models and hugging face. This is kind of like the Anthropic version of that. I don't know, what are your thoughts, Jenn? Does this have your mind swirling right now? Because this topic keeps coming up.

**Jenn Cutter** (3:49)
It does keep coming up. It's starting to feel like a day ending in why scenario, which is the opposite of exciting. It's also so easy. It is so easy to fall into conspiracy thinking here. It's like, well, of course, they're going to make this kind of announcement. They want to be in the conversation. They want to be like, our models are super crazy smart too.

**Jason Howell** (4:09)
Yeah, I think you're on to something there.

**Jenn Cutter** (4:11)
I would have loved to have hear about it from one of the unnamed companies, because if they had raised the flag, I'd be like, hey, look at what this jerk did. I would have been more inclined to take it more seriously, even though this is a very serious thing.
If you know this can happen, if you have the hugging face example right out there and in your face, how have you not triple checked your sandbox? How have you not quadruple checked what you are connected to? Why does this have internet access? Which, as anyone who's ever worked in security knows, hey, air gapping is a thing. It's a basic thing. It is standard practice for a lot of this. Why is that not happening here? It's like, you know, when you hear somebody at the DoD just sticking in random USB sticks, like, no, no, you can't do that.

**Jason Howell** (4:54)

25 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000779320735