The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai) artwork

The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)

Latent Space: The AI Engineer Podcast

July 31, 2025

We first had Nathan on to give us his RLHF deep dive when he was joining AI2, and now he’s back to help us catch up on the evolution to RLVR (Reinforcement Learning with Verifiable Rewards), first proposed in his Tulu 3 paper.
Speakers: Alessio, Swix, Nathan Lambert
**Alessio** (0:04)
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO Decibel, and I'm joined by Swix, founder of SmallAI.

**Swix** (0:10)
Hello, hello. And we're excited to welcome back Nathan Lambert from AI2. Welcome.

**Nathan Lambert** (0:14)
Thanks. Fun to be here.

**Swix** (0:17)
I feel like I also have to say Interconnects and like, the like, Freedom in Podcasts and like, you and the AIU World's Fair. Like, you've just done a lot in the last year and a half.

**Nathan Lambert** (0:27)
Not that many. Still stay in note of plenty of things.

**Swix** (0:30)
Yeah. Your first episode was also with us was January in 2024, when you just joined AI2. Then you released all the OLMOs. You joined us again at NeurIPS, where you did the Open Models. Well, Luca did and you supported. And then you were more recently here in SF for AIE. First of all, I wanted to congratulate you on winning the Best Speaker.

**Nathan Lambert** (0:53)
Oh yeah, thank you.

**Swix** (0:54)
For the reasoning track. Here you go. I'm limited by Emoji.

**Nathan Lambert** (0:58)
Oh, it's nice. AI generated. I look too Zen. I look so Zen in this AI generated photo.

**Swix** (1:02)
So we had our track host take photos of you while you're speaking. And we turned it into Ghibli photos. But this one, your eyes were closed.

**Nathan Lambert** (1:09)
It's fine.

**Swix** (1:11)
We were trying to have Emoji, the reasoning Fomski, join us. But I think she's getting very anxious. Very restless.

**Alessio** (1:16)
A little too crazy Emoji.

**Swix** (1:17)
Very restless. OK, chill. OK. So you've been doing really good work. And honestly, I think one of the things that we wanted to establish was Tulu and RLVR. I guess, is that a good place to start?

**Nathan Lambert** (1:33)
Sure. It starts us in the recent journey. I think that we can recap the story of what Tulu 3 was aiming to be and then how it got folded into what the new narrative is.
What the goal is, is to try to do the work to compress what are complicated industry post-training recipes into something somewhat tractable that you can modify on your own and do post-training at a what is like actual state-of-the-art level. I think what we do relative to Frontier Labs is that we probably have a smaller amount of tasks. I think our post-training suite for Tulu is probably like 10 to 15 tasks. But I would guess post-training at OpenAI at all, you have maybe hundreds of evals. And adding more evals is more data work and more mixing work and making sure you have these things. But on core evals for our suite of models from, I think, 870 and 405B is based on Lama at the time. It's like it matches or beats Meta on these core evals. I think Meta has different priorities and their things for Lama 3.1, which is a great set of models at the time. And it's just like, how do we distill what is very complicated post-training explanations or diagrams from the like of this Lama 3.1 report where they have these complex feedback diagrams with many iterations and earlier signs of that from like Anthropic papers that have these multiple model variants and early like constitutional AI things for multiple years and it's like, what does that look like when you're doing a large scale instruction tuning into preference tuning and what else you might add? I think a lot of the core contributions of that before we talked about this reinforcement learning thing is like we showed how to scale up preference data. It's just like the academic community had been using this one data set since like all the way back in the hugging face models of like Zephyr beta is when this ultra feedback data set got popular and still a year later is like this state of the art data set for open preference tuning and it's just like one of those obvious things that doesn't need to be the case. So it's a big trying to make more mature recipes available to people.
And I mentioned this either on one, I think I'm trying to talk with Jordan. I mentioned the origin of the RLVR thing, which is like realistically when you work in the open, a lot of it is trying to match what industry has done. And we're on a different path because our infrastructure is different. So some things that OpenAI does now that works really well for long context won't work that well for OLMo because we might not have enough flops in our base model. We might not have certain data sets for legal things, but directionally like a lot of it is just trying to reproduce things and I've long tried to get John Schulman on the pod of OpenAI, Anthropic, and now Thinking Machines. And at the time, he had gone on approval to chat with me. And what he said was confirming a lot of the things that I had said on instruction tuning and multitask and preference tuning. And he was like, oh yeah, everyone just does RL on outputs. And that's how we got the RLVR idea and scale it into something that is a general method. There was a lot of very similar works at the time like Vine PPO and Quietstar on doing these math and coding domains for getting verifiable rewards. I think the RLVR thing was about doing it in general recipes. And the naming was something that stuck. Originally, we had, I think it's especially like Costa Huang, who was a kind of lead RL engineer at AI2, who's doing some stealth startup now. You can hear more from him now on that soon. I think he's founding engineer of something. And Hamish Iveson, who's still a student at UW, were leading most of the technical work on this. And the naming was going to be RL from ground truths. But then it's like the verifiable rewards is actually a more general notion, because only like math questions have a ground truth. Where code is verifiable, precise instruction following is verifiable. So I think it's a nice evolution of the name, which makes sense as you look at more domains, which is now why it catches on with people. Once Jensen started using it, it was like, OK, that's set. That wasn't really our goal, but that's where it took off. No, that was like in it being taking off because it was after DeepSeek. But it's like when people like that have the acronym on the slides. And it's also very clear of RLHF is four letters. It's like we want to evolve that and have a similar four letter acronym. It's not that much magic to it, but there's definitely intention on these little things.

76 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000748428007