Fake AI Safety Experts Are Stealing Real Cyber Contracts | Brian Chau artwork

Fake AI Safety Experts Are Stealing Real Cyber Contracts | Brian Chau

MTS

September 15, 2026

Brian Chau joins MTS to discuss his investigation into Irregular's role in recent AI cybersecurity incidents, the involvement of effective altruism organizations in AI evaluation firms, and the need for traditional cybersecurity expertise in AI testing. Turn ideas into software people love.

Speakers Brian Chau

TopicsNews

Brian Chau (0:00)

I reported last week on these firms that are all paid by one of two foundations. Over the course of one Guardian article, this was an article that was published by The Guardian, which was funded by Open Philanthropy, was written by a journalist whose salary was being paid by Open Philanthropy, described a project, the AI Futures Project, which was funded by Open Philanthropy, and quoted three fake experts who are all funded by Open Philanthropy or by Survival and Flourishing Fund, either directly or indirectly.

This was the construction of the entire Potemkin field.

SPEAKER_2 (0:38)

And we're back, we're live with Brian Chau, who's the founder of Effort.News, Effort.News, which is an independent, open-source investigative journalism project that has published a lot of interesting things recently. Brian, welcome to MTS.

Welcome back.

Brian Chau (0:54)

Great to be back.

SPEAKER_2 (0:55)

Yeah, we had you on on launch day. You just posted a story like 10 minutes ago about Irregular, which is the company that was doing, I believe, evals for opening iAnthropic and Meta that led to the models hacking into things. So tell us about Irregular and what you learned.

Brian Chau (1:17)

Yeah. So there's been the story that's been very common, that AI agents went rogue, just decided to start hacking things by themselves, created new civilizations. And I'm here to tell you, all of that is bullshit. Because what actually happened was that three AI companies that we know of that had these security problems contracted with Irregular, they instructed the AI models to exploit certain vulnerabilities. And then they did exploit those vulnerabilities using the Internet, which Irregular gave them. They gave them this in these models Internet access, they say accidentally, and very clearly instructed the models to engage in these cyber attacks.

So the actual story of what happens is much plainer when it comes to liability. And much plainer when it comes to what the policy implications should be. There's this one firm that has had these vulnerabilities, they are the common factor with all three of these companies in terms of creating these environments, which were faulty, which allowed models to access the Internet, and which in the course of asking the models to attempt to exploit these specific vulnerabilities, led to them hacking into real targets.

So the big takeaway here is that, number one, there's this key security flaw in this one firm that has led to all these headlines going across the entire internet, and now beginning to influence US policy. And that the firm itself is closely linked to effective altruism Israel. In fact, one of the co-founders is literally on the board of effective altruism Israel.

SPEAKER_2 (3:17)

Interesting. So what are the implications of irregular being linked to EA? Seems like much of the AI world is linked to EA in some form.

It would be like difficult to find a company or a person that isn't, like at most to connections away from some EA org.

Brian Chau (3:36)

So the main issue here is the narrative that they're now spreading. Of course, we wouldn't expect any company to go out and say, hey guys, we had severe flaws in our sandboxing. We actually completely fucked up. We asked models to hack your company and those models hacked your company. And actually, we should be put in jail for that. No one's going to go out and say that. But there's an interesting combination here. Some people ask, is the real reason that people spread these narratives for self-interest or because of ideology? And this is a scenario where it's very hard to pull the two apart.

The effective altruism connection is important because that's where a lot of this idea of misalignment and the rogue agent theory, that's where it comes from.

Anthropics own logs with the Anthropic Irregular Incidents, one of them, found that this was completely actually, that when they instructed the models not to access the internet, the models did not access the internet.

Here's an exact quote. None of the prompts, these are the prompts involving the incident. None of the prompts stated which systems were in scope for that exercise or constrained where Claude could search for the flag. So we had the scenario where not only could you stop the models from hacking into the real world targets by saying, by stopping them from having internet access altogether, but you could stop them from hacking into the real world targets by telling them not to hack into the real world targets. In other words, irregular and entropic are entirely responsible for the hacks that these models did. It was not misaligned. It was in fact completely aligned with the prompts and with the environment that was provided to them.

14 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Fetch the whole transcript

The demo key returns a sample episode in full, no card needed:

request
curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

request
curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000789788611