**Nathan Lambert** (0:00)
It's probably not worth the effort to spend all your time on preference tuning when you could just be making better data and better pipelines, which is what 2.3 is about. Inspired by the transition we're seeing with the LlAMA report, with Trapbot Arena, it's turned into a hockey stick again, where we have these incremental scores, and the OpenAI and Google are skyrocketing their scores. And the philosophy is, how do we try to understand what the open groups should be doing, and where there are hills to climb when you're increasing the complexity substantially in post-training? If you can have humans and LLMs do preference data, which do you send to humans vs. LLMs? And that, I think, solves a lot of the problems, which is like, there are definitely things that we want humans giving the answer on, but there are a lot of mechanical tasks that we can outsource to LLMs. There's a lot more in post-training that is not really touched, and I think that the opportunity is high. Because the big thing is, how do you develop character? Character is something that you don't have evaluation for in your models, and our models, if you compare them to Claude, will not have as consistent of a character.
**Erik Torenberg** (1:00)
Regardless of how you felt about the outcome of the election, I think we were all united in looking forward to an end to the constant fundraising emails and text messages. Unfortunately for me, they haven't stopped, even now more than a week after the election. And that's to say nothing of the normal commercial spam coming at me from all directions at all times. It turns out that most of this noise is caused by data brokers. These companies aren't just collecting your contact details, they're gathering everything from your social security number and financial records to your online shopping habits. And they're now working with insurance companies, which could potentially impact your rates. That's why I'm excited to now be using Incogni. Incogni contacts five types of data brokers. Marketing, recruiting, financial information, risk mitigation, and people search sites, and demands that they remove your information. Then they continue monitoring on your behalf to prevent data recollection. Take your personal data back with Incogni. Or protect up to four family members with their family and friends plan. They offer a 30-day money-back guarantee if you're not satisfied. So take your personal data back with Incogni. Use code REVOLUTION at the link below and get 60% off an annual plan. That's incogni.com/revolution.
Hello, and welcome back to The Cognitive Revolution. Today, my guest is Nathan Lambert, author of the popular Interconnects newsletter and machine learning researcher at the Allen Institute for AI, which today is releasing Tulu 3, one of the most comprehensive open-source efforts to diffuse the understanding and practice of frontier post-training techniques for large language models that we have seen to date. By systematically working to match Meta's post-training performance using the same LlAMA base model and sharing all of their findings and data publicly, Nathan and the team at the Allen Institute have illuminated what has historically been one of the most opaque aspects of large language model development. And this conversation represents one of the most detailed discussions of this topic that you can find anywhere online today. We cover the full spectrum of post-training techniques, including supervised fine-tuning, multiple flavors of preference-based reinforcement learning, and a new technique called reinforcement learning from verifiable reward, which rewards the model for accurately answering questions with objectively correct ground truth answers. At each step, we dig into the practical details that make these techniques work, the associated compute requirements, data generation strategies, and the value derived from each, as well as the experimental designs that are used to measure performance while exploring the vast space of possible training recipes. We even explore some fascinating emergent behaviors that echo the frontier reasoning capabilities that we've recently seen from OpenAI's O1. Nathan's frank discussion of both the technical and organizational challenges of this work, including in a few moments where he acknowledges aspects that are not yet well understood, provides a super useful window into what it takes to develop state-of-the-art models. And the fact that they ultimately succeeded in matching LlAMA performance with a team of just 10 to 15 people makes their approach one to study closely. Now, how long this level of open development can continue into future generations of models remains, in my mind, an open question. Even with billionaire estate backing, human-generated preference data and annotations are cost-prohibitive for the Allen Institute, and the synthetic data generation techniques used in this project may or may not be available going forward if frontier developers follow OpenAI's lead and choose not to release their O1-style reasoning traces. That said, Nathan expects that the community will ultimately figure something out, and just before publishing, we've seen Chinese AGI company DeepSeq announce a new O1-style model called DeepThink, which does seem to show its work and is reportedly going to be open-sourced in the near future. A development that could shift the question from whether or not such open reinforcement learning projects can continue to whether or not they should. And which should definitely cause anyone who thinks that Western companies can maintain a comfortable lead over Chinese companies from here all the way to AGI to stop and rethink their assumptions. In any case, this episode is one of the highest leverage pieces of content we've produced on this show to date. It's full of concrete, practical insights that were previously scattered throughout the literature or simply hidden behind closed doors. And I am really grateful to Nathan for such an open and technically detailed conversation. If you're finding value in the show, we appreciate it when listeners take a moment to share it online, post a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. And your feedback is always welcome, via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. With that, I hope you enjoy this super deep dive into the frontiers of large language model post training with Nathan Lambert of the Allen Institute for AI. Nathan Lambert from the Allen Institute, creator of Tulu. Welcome to the Cognitive Revolution.
98 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000677801656