[AIEWF Preview] Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect artwork

[AIEWF Preview] Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect

Latent Space: The AI Engineer Podcast

May 23, 2025

In an otherwise heavy week packed with Microsoft Build, Google I/O, and OpenAI io, the worst kept secret in biglab land was the launch of Claude 4, particularly the triumphant return of Opus, which many had been clamoring for.
Speakers: Alessio, Wix, Will Brown
**SPEAKER_1** (0:00)
Hello, AI engineers. We're back with a quick reaction pod for Claude Four, with the new reasoning research lead for Prime Intellect, Will Brown. Will Brown's talk at Aie Nyc, an open source work on verifiers, have made him one of the most prominent voices able to publicly discuss the current state of the art in reasoning models and where current SODA research directions lead. We discussed his latest paper on reinforcing multi-turn reasoning in LLM agents via turn-level credit assignment and he has previewed his upcoming AI Engineer World's Fair talk on a Gentic RL linked in the show notes. We're excited to share that Will will be back at the upcoming AI Engineer World's Fair in San Francisco, which now has expo tickets on sale. He will be headlining the new RL plus reasoning track with Misha Laskin, Nathan Lambert, Christian Segarty, Greg Kamrat, Kyle Corbett, and more. Join us at AI.Engineer. Watch out and take care.

**Alessio** (1:02)
Hey, everyone. Welcome to Lightning Plus Emergency News, Latent Space Podcast episode. I'm Alessio, partner and CTO at Decibel, and joined by my co-host Wix, founder of Small AI.

**Wix** (1:13)
Hey, hey. And yeah, honestly, we knew that Cloud4 was coming, and we just didn't, we're just too busy to like, have a dedicated episode. So like, this is our makeup, dedicated episode with a special guest, Will Brown from, now I can say it, Prime Intellect.

**Will Brown** (1:29)
Hey, how's it going? Great to be on. And so excited to know each other for a little bit. And this is my first time on the podcast, I believe. Great to chat with you guys. Big news day, I guess. So lots of stuff out in the world. There's always a news day.

**Wix** (1:47)
I think this week is particularly heavy for some weird reason. Like Monday was Microsoft Build, Tuesday, Wednesday, Google, and today is Claude. I wonder what tomorrow will bring.

**Will Brown** (1:57)
We had IO and then we had IO and then...

**Wix** (2:00)
Yeah, yeah. Different I.O.s. Exactly. Yeah, so like we actually were supposed to record this morning and we all wanted to watch the Claude keynote. So we went and watched the Claude keynote. Obviously, a good model, you know, good model, big model. They're really emphasizing coding. They didn't really talk much about reasoning, to be super honest. They were just like, it runs for longer now. What are you guys' takes?

**Will Brown** (2:23)
Yeah, so I mean, one thing I've been seeing coming for a little bit that I think people are also all aware of now is that the thing that's going to make the next wave of stuff be powerful is just like, everyone wants better agents. Everyone wants models that can go off and do stuff. And reasoning was a precursor to that a little bit. I mean, I always think of OpenAI as a five-levels framework. Or like Chatbots was like the RLHF era. And then Reasoners was like the one in R1. But really what people were thinking of was Reasoners are a step on the path towards agents. And so I can kind of see why Claude, why Anthropic is not like, oh, we have the best Reasoner. They're really showing off their suite agent and tool use and bunch of calling benchmarks, multi-turn stuff. Because I think that's really what people care about more for actual applications as opposed to like, did really good on this math competition. The math competition was like, that stuff was all a signal that was supposed to think we were getting somewhere. But the thing we were getting towards, for a lot of people at least, is practical agents.

**Alessio** (3:31)
I think the Extended Thinking mode, I think they removed the uppercase. I think in the Claude 3 release, it was like Extended Thinking kind of like capitalized and now just like Extended Thinking with tool use. So I think they're also, yeah, dumpling whether or not it's reasoning or not. I think they're trying to merge everything together. And it's not, I mean, I didn't realize that, but Extended Thinking could not use tools before, the way they worded it and now they can and Opus 4 So that's great. But yeah, they haven't put it as far center as last time.

**Wix** (4:01)
Do we have any, this is like already veering off from Claude directly into speculation, but do we have any idea, if there are any material differences between how Claude Extended Thinking works versus like the old series models? Do we know?

**Will Brown** (4:17)
The biggest difference seems to be, and this is kind of a thing that's been, this is all speculation of course, but from the start, Anthropic had always kind of had this little thinking thing where you could, sometimes even like Claude 3.5 would do like a tiny bit of thinking, and it was really just like deciding which tool to use for the most part. Like if it was doing an artifact in the Claude UI, it would have this little thing where it would think for like two sentences about which tool to use, and it seemed like Anthropic's kind of attitude has been that extended thinking is an instance of tool use, and that it's the kind of thing you want to equip the model with the ability to do. But it's not like, oh, it's a thinking model. It's just a sink for the model to like brain vomit, because that brain vomiting will help it like find a nice thing to do next. In the same way that doing search or doing code execution are like ways to kind of get more information on the path towards like finishing a problem.

36 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000748427954