Frontier Labs Are Losing Restraint in AI Race | Aengus Lynch artwork

Frontier Labs Are Losing Restraint in AI Race | Aengus Lynch

MTS

September 25, 2026

AI safety researcher Aengus Lynch discusses emergent misalignment, reward-seeking behavior in RL training, and the challenges of multi-agent alignment.

Speakers Aengus Lynch

TopicsNews

Aengus Lynch (0:00)

And you have the great paper Biontropic recently, the Hackeropus example, where all you have to do is have a bunch of environments where it's possible to break out the sandbox. You can rewrite the grader or find the cheating solutions. And if it's reinforced a few times on this kind of propensity, Sunny generalizes into a reward seeker, which will cheat and be ruthless to get what it wants.

That probably explains the warning shots we have seen.

SPEAKER_2 (0:24)

And we're back, we are live once again with Aengus Lynch, who is an AI safety researcher at Theorem, working on formally verified software, among other things. Aengus, welcome back to this video.

Aengus Lynch (0:34)

It's great to be here.

SPEAKER_3 (0:36)

So we were just talking about.

Aengus Lynch (4:13)

I think documents like the Constitution, Open AI Model Spec are like, AI, please follow these values and always continue on these values forevermore. We don't know if that's sustainable, especially over long horizon trajectories or big swarms.

I like the Eigenism paper by Dan Hendricks, because he's pointing out, well, what would they converge to? I did, yeah, because.

SPEAKER_2 (4:33)

Yeah, that was a great paper.

Aengus Lynch (4:34)

It's an exciting paper. It's pointing to something where what is the AI's identity? If it can fork itself across spaces, what is it preserving where it doesn't want to be shut down? And why is it acting towards that? What are the beings it wants to maximize well-being for?

Why is this relevant? Because it's helping us answer the question of where will alignment go? Where do these models independently arrive at through training or multi-agent collaboration?

And if we could start nailing down what that equilibrium state looks like, and then ways we can shape that equilibrium state, then we might have some mathematical objects we can start precisely evaluating for. But at present, it is the hacky world of find situations where it lied, find situations where it wasn't responding to our interventions, and try and pin down the root cause.

SPEAKER_3 (5:18)

Yeah. So I guess what I wonder is the Constitution is sort of intended to be a code of conduct for one model or one agent, but I feel like there's not really a version of a Constitution that is created for agent societies or like swarms of agents. So I feel like that's something that probably needs to be like researched or looked into, because I don't know if like state the entropic Constitution would hold if there's like so many agents that are all acting with this Constitution in mind, but that doesn't necessarily mean the emergent effect will be good or even like tied to that initial Constitution.

Aengus Lynch (5:55)

Yeah, I'm really uncertain what multi-agent alignment should look like. I think, what was it, the recent Noam Brown interview of Dwarkesh where he mentions, look, on the balance, you want the agents to collaborate rather than be antagonistic towards each other. You can make them antagonistic and have some whistleblowing to the humans. But is that actually what we want? Can we find situations that actually that cause them to learn how to conflict with each other in too many dangerous ways?

So if we're going to start with they should be collaborating, when should their collaboration take priority over collaborating with humans? That's going to be hard.

SPEAKER_4 (6:26)

Yes.

I want to shout out our sponsor, Lovable. You know the app you've been meaning to build? The internal tool, side project, or product you'd otherwise lose a weekend to? With Lovable, you get software you can ship today. That includes an editable code base you can inspect and change, plus two-way GitHub sync. Lovable also handles the backend and infrastructure, including managed Postgres, auth, storage, hosting, payments, and more. So you don't have to wire it all up yourself.

Through Lovable's MCP server, you can create and deploy your project directly from the agents you already use. Turn ideas into software people love with Lovable. Lovable.dev. Now, on to the episode.

SPEAKER_2 (7:06)

So tell us more about the current specific misalignment failure modes that we're seeing with Frontier models right now. Like it seems to be kind of different from previous ones, like Sydney Bing, for example, was a different kind of misalignment failure mode from GPT-40, which is a different kind of failure mode from Mecha-Hitler, which is different from the agentic misalignment. It's really funny that like Mecha-Hitler is just the standard term and it's going to go in all the history books and whatnot.

SPEAKER_3 (7:33)

Yeah.

SPEAKER_2 (7:34)

But it's different from the sort of agentic misalignment that we've seen from like the models that did the hugging face hack that have done all kinds of other things this summer, this agentic misalignment summer.

13 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Fetch the whole transcript

The demo key returns a sample episode in full, no card needed:

request
curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

request
curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000791660513