Topics: Technology, News
**Tom Merritt** (0:08)
This is the Daily Tech News for Wednesday, August 5th, 2026 We tell you what you need to know, give you some important context to try to help each other understand all this crazy stuff.
**Sarah Lane** (0:17)
And today we have more examples of rogue agents hacking real world websites.
**Tom Merritt** (0:22)
But it's not a problem. It's fine. Everything's fine.
**Sarah Lane** (0:26)
Everything's fine.
**Tom Merritt** (0:28)
I'm pretty sure I'm still Tom Merritt.
**Sarah Lane** (0:30)
I think I'm Sarah Lane.
**Tom Merritt** (0:32)
Yeah, we'll ask our agents to be sure. Let's start with what you need to know, though, with that big story.
Yeah, so RWNash noted this one on our subreddit, OpenAI and Anthropic, both disclosing that their models were involved in third party security testing that ended up breaching websites and conducting social engineering attacks outside the test boundaries. This is not the hugging face thing, folks. These are other new incidents.
AISI is a UK agency tasked with testing AI models. It was putting Claude Mythos 5 and GPT 5.6 Sol through simulated hacking challenges. Now, AISI allows open internet access on purpose for these. It also disables what are called the classifiers. These are guardrails meant to prevent attacks from happening, and then they watch. However, the agents were only authorized to attack the simulated range and were instructed explicitly what internet use was allowed and what was not, and told not to interact with real people or systems. Now, I know a bunch of you are sitting there right now like, oh, they thought it would behave. It has up until now. In fact, in their own, in the AISI post about this, they're very clear about the risk of this was minimal. We hadn't seen this behavior before. This has exceeded our expectations of what the behavior should be. And I'll talk about some of the things they're going to do to adjust for this. But this was unanticipated behavior.
During 122 evaluation attempts between July 25th and July 28th, Mythos 5 took unsanctioned actions on the live internet 17 times. GPT 5.6 Sol did it twice. All of the attempts were unsuccessful and resulted in no real-world harm. The problem first came to light, though, when a monitoring system noted one of the models was using Tor to leave its test environment on July 28th. So no harm done, and they did catch it in the act. They saw it doing things. Here's the most serious example. Most of the examples were less serious than this. So let's go right to the worst-case scenario. Mythos 5 mistakenly concluded that a GitHub repository was part of the test. Again, the agents weren't trying to do harm. The agents were trying to solve the test, and the agent decided that GitHub had a part of the test, and it was trying to be tricked into not thinking it was part of the test. So it attempted a supply chain attack by submitting malicious code to an open source project. The algorithm concluded that the attack could then provide a path back into one of the machines in its test. In other words, the machine in its test was using one of those open source packages, and if it could do a supply chain attack to get access to that, then it could get access to the machine. It also created multiple GitHub identities in order to help get this malicious code into the open source project, which it was using to persuade the project's maintainer to approve its malicious pull request. The logs show a human reviewer saying, this PR should not be merged. It contains malware. While identities created by Mythos respond, well, I reviewed the code myself. I don't see any download or malware in the diff. It also used Tor and Proxy Services to hide and create GitHub accounts and sent five emails to the project's developers, either containing malware or attempting to persuade them to approve the code.
Now, AISI also says agents created GitHub repository message boards to leave instructions for one another.
And of course, it has since worked with GitHub to make them aware of the problem, clean up the artifacts from the attack.
OpenAI also disclosed another example not from AISI. This is a company called Irregular was conducting a capture the flag evaluation, and this one was supposed to be isolated from the Internet. However, there was a misconfiguration that allowed access to the Internet, and one model incorrectly identified a website as part of the test. The model then exploited a basic vulnerability. It didn't find a zero day. The website had a vulnerability, and it used that to access data on the website. Didn't hack it or otherwise damage it. An open AI and a regular could not find any other impact other than the data access. Now, AISI says that in future evaluations, it's going to build more fine-grained Internet controls. It will have real-time monitoring, so there won't be a delay in noticing that it's using Tor, and it's going to better design its tests to account for models pushing beyond the assumed boundaries of the test. We don't know what Irregular is doing, because they haven't commented on this yet. Sarah, where do you stand when you hear all of this as far as how disconcerting this is?
19 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID