**Dwarkesh Patel** (0:00)
Okay, I'm joined again by my friends, Sholto Bricken, wait, fuck.
**SPEAKER_2** (0:06)
Did I do this last year? No, no, no, you named us differently, but we didn't have Sholto Bricken and Trenton Douglas.
**Dwarkesh Patel** (0:14)
Sholto Douglas and Trenton Bricken, who are now both at Anthropic. Yeah, let's go.
Sholto is scaling RL, Trenton's still working on mechanistic interoperability. Welcome back.
**Sholto Douglas** (0:29)
Happy to be here.
**Trenton Bricken** (0:30)
Yeah, it's fun.
**Dwarkesh Patel** (0:31)
What's changed since last year? We talked basically this month in 2024, now we're in 2025, what's happened?
**Sholto Douglas** (0:37)
Okay, so I think the biggest thing that's changed is RL and language models has finally worked. And this is manifested in, we finally have proof of an algorithm that can give us expert human reliability and performance given the right feedback loop. And so I think this is only really being conclusively demonstrated in competitive programming and math, basically.
And so if you think of these two axes, one is the intellectual complexity of the task, and the other is the time horizon of which the task is being completed on. And I think we have proof that we can reach the peaks of intellectual complexity along many dimensions. But we haven't yet demonstrated long-running, agentic performance. And you're seeing the first stumbling steps of that now, and should see much more conclusive evidence of that basically by the end of the year. With real software engineering agents doing real work. And I think Trenton, you're experimenting with this at the moment.
**Trenton Bricken** (1:30)
Yeah, absolutely. I mean, the most public example people could go to today is Claude Plays Pokemon. And seeing it struggle in a way that's kind of painful to watch, but each model generation gets further through the game. And it seems more like a limitation of it being able to use a memory system than anything else.
**Dwarkesh Patel** (1:51)
I wish we had recorded predictions last year. We definitely should this year.
**Trenton Bricken** (1:55)
Yeah, hold us accountable.
**Dwarkesh Patel** (1:56)
That's right. Would you have said that agents would be only this powerful as of last year?
**Sholto Douglas** (2:01)
I think this is roughly on track for where I expected with software engineering. I think I expected them to be a little bit better at computer use. But I understand all the reasons for why that is, and I think that's well on track to be solved. It's just a sort of temporary lapse. And holding me accountable for my predictions next year, I really do think end of this year, or this time next year, we have software engineering agents that can do close to a day's worth of work. For a junior engineer. Or a couple of hours of quite competent and independent work.
**Trenton Bricken** (2:35)
Yeah, that seems right to me. I think the distribution is pretty wonky, though. For some tasks, I don't know, like boilerplate, website code, these sorts of things.
It can bang it out and save you a whole day.
**Sholto Douglas** (2:45)
Yeah, exactly.
**Trenton Bricken** (2:47)
Yeah, I think that's right.
**Dwarkesh Patel** (2:47)
I think last year you said that the thing that was holding them back was the extra nines of reliability. I don't know if that's the way you would still describe the way in which these software agents aren't able to do a full day of work, but are able to help you out with a couple of minutes. Is it the extra nines that's really stopping you, or is it something else?
**Sholto Douglas** (3:04)
Yeah, I think my description there was in retrospect probably not what's limiting them. I think what we're seeing now is closer to lack of context, lack of ability to do complex, very multi-file changes, and maybe scope of the change, or scope of the task in some respects. They can cope with high intellectual complexity in a focused context with a scoped problem, but when something is a bit more amorphous, it requires a lot of discovery and iteration with the environment, this kind of stuff, they struggle more.
So maybe the way I would define it now is the thing that's holding them back is, if you can give it a good feedback loop for the thing that you want it to do, then it's pretty good at it. If you can't, then they struggle a bit.
**Dwarkesh Patel** (3:56)
For the audience, can you say more about what you mean by this feedback loop if they're not aware of what's happening in RL and so forth?
**Sholto Douglas** (4:01)
Yes. The big thing that really worked over the last year is maybe broadly the domain is called RL from verifiable rewards or something like this, where a clean rewards.
136 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000709488338