Both Sides Are Overconfident On AI | Toby Ord artwork

Both Sides Are Overconfident On AI | Toby Ord

MTS

September 29, 2026

Oxford researcher Toby Ord breaks down the power-law mathematics of agent swarms, how parallel AI coordination alters recursive self-improvement timelines, and why transformative AI could arrive before 2030. Turn ideas into software people love.

Speakers Toby Ord

TopicsNews

Toby Ord (0:00)

Dario has this thing about a country of genuses and a data center. You pay a penalty for parallelizing things in terms of the amount of compute you need to use, but the advantage that you gain comes in speed. If you had 100 times as many agents, you'd solve it in a 10th as much time, but you'd pay 10 times as much for it.

SPEAKER_2 (0:20)

Senior researcher at Oxford's AI Governance Initiative. He's the author of The Precipice, a book about existential risk and the future of humanity. He recently wrote an article on swarm scaling. He's been talking about all kinds of AI scaling axes recently.

Toby, welcome to MTS.

Toby Ord (0:37)

It's great to be here.

SPEAKER_2 (0:40)

Tell us about your new piece on swarm scaling. Explain the thesis for our audience, like how to capability scale when agent swarms increase in size.

Toby Ord (0:49)

Yeah, so we've seen a bunch of agent swarms, ranging from a handful of agents working together to solve a task for a user, through to 1,200 agents that were meant to be working alone, but found each other on open AI servers and worked together to cheat on their tests and ultimately hack Hugging Face. Then even 10,000 different agents working together to solve maths problems, including this Navier-Stokes problem. There's a natural question about what do you get for having a thousand or 10,000 agents working together? How much smarter do they get? What I thought would be interesting is to look at the normal scaling laws that we have for, if you take a single agent and you give it longer and longer to solve a problem. If you give it 100,000 tokens or you 10x that or you 10x it again, what you tend to find is that there's a fairly steady improvement in capabilities every time you increase the amount of tokens by some factor. I wanted to try to see, is there a way of converting things together? To say, is it the case that if you can 10x the amount of compute, are you better off putting that into a longer chain of thought or are you better off putting it into 10 times as many agents? How does that work? What I found, at least what I hope I've found, is that there is a way of converting them, that there's a way, this kind of power law relationship between them.

What you get is that you can take the number of tokens that you want to scale up by, you can raise it to some power, some special kind of exponent, lambda, and what you get back is the equivalent as how many more agents you would need. Actually, it's the other way around. If you scale up the number of agents, let's say by a factor of 100, and then you raise that to this special power, let's say a half, you get 100 to the power of a half, which is 10, and so 100 times as many agents gives you about the same as if you had 10 times as long a run.

That's the best I can put it on an audio stream.

SPEAKER_2 (3:27)

So what does this imply for the rate of progress of RSI? Obviously, Lambda plays into it tremendously, but how specifically does swarm scaling accelerate things? I imagine some RSI tasks are much more parallelizable than others.

Toby Ord (3:42)

Yeah, so this parameter Lambda, which I said was like a half in my example, that's one of the numbers I got out of OpenAI's data when I tried to back out, what would this number be? Also, it could be a little bit higher than that. If it was as high as one, that would mean that they scale perfectly. That's like the kind of idea of the mythical man month, you know, where if you have a project that takes like, you know, one person, 100 months, maybe 100 people could do it in one month. Things tend not to be as parallelizable as that in reality.

But actually one of the results from the system card of Opus 5.5 just came out, they have some graphs as well, which you can try to estimate lambda. And in one of them, lambda is about one, which means it's roughly speaking perfectly parallelizable. That particular task is kind of, if you look at it, you kind of, you can see why it's parallelizable. The task is to create a knowledge base. So you're given, I think, a whole lot of code and you've got to try to like write up some documentation for it, such that then if a human is given some questions, they can use your documentation to answer the questions correctly. And what happened was that this Opus 5.5 agent kind of created 100 different subagents and just parceled out all the different kind of parts of the knowledge base between the different agents. Everyone just worked on their own bit and didn't have to really talk to each other. So that was like a really simple kind of case where you could get the scaling to be really good, basically as good as if you just gave an agent longer and longer time. But in some other cases, the agents get in each other's way a lot. Economists call this stepping on toes, and it's their cute little name for it.

17 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Fetch the whole transcript

The demo key returns a sample episode in full, no card needed:

request
curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

request
curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000792114745