**Alessio** (0:05)
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO of Residence and Decibel Partners, and I'm joined by my co-host, Swix, founder of Small AI.
**Swyx** (0:15)
Hey, and today we have Dr. Nathan Lambert in the house. Welcome.
**Nathan Lambert** (0:18)
Thanks guys.
**Swyx** (0:18)
You didn't have to come too far, you got your PhD in Berkeley, and it seems like you've lived there most of the time in recent years. You worked on robotics and model-based reinforcement learning on your PhD, and you also interned at FAIR and DeepMind. You bootstrapped the RLHF team at HuggingFace, and you recently joined the Allen Institute as a research scientist.
So that's your quick bio. What should people know about you that maybe is not super obvious about you on your LinkedIn?
**Nathan Lambert** (0:43)
I stay sane in various insane sport and ultra-endurance sport activities that I do.
**Swyx** (0:50)
What's a ultra-endurance sport activity?
**Nathan Lambert** (0:52)
Long distance trail running or gravel biking. Try to unplug sometimes, although it's harder these days.
**Swyx** (0:59)
Yeah, well, just the Bay Area is just really good for that stuff, right?
**Nathan Lambert** (1:02)
Oh yeah, you can't beat it. I have a trailhead like 1.2 miles from my house, which is pretty unmatchable in any other urban area.
**Swyx** (1:10)
Oh, pretty excellent. You also have an incredible blog, Interconnects, which I'm a fan of.
And I also recently discovered that you have a new podcast, Retort.
**Nathan Lambert** (1:20)
Yeah, we do. I've been writing for a while, and I feel like I've finally started to write things that are understandable and fun after a few years lost in the wilderness. If you ask some of my friends that I made read the earlier blogs, they're like, oh, this is yikes, but it's coming along, and the podcast is with my friend Tom, and we just riff on what's actually happening on AI and not really do news recaps, but just what it all means and have a more critical perspective on the things that really are funny, but still very serious happening in the world of machine learning.
**Swyx** (1:52)
Yeah, awesome. For people who are new to your work, what would you highlight as your greatest hits so far on interconnects, at least?
**Nathan Lambert** (1:59)
So the ones that are most popular are like timely and or opinion pieces. So the first real breakout piece was in April when I also just wrote down the thing that everyone in AI was feeling, which is like, we're all feeling stressed that we're going to get scooped and that we're overworked, which is like behind the curtain, what it feels to work in AI. And then a similar one, which we might touch on later in this was about my recent job search, which wasn't the first time I wrote a job search post, but people always love that stuff. I mean, it's like easy for me to do in a way that it vary on brand and it's very helpful. Like I understand that until you've done it, it's hard to share this information. And then the other popular ones are various model training techniques or fine tuning. There's an early one on RLHF, which is, this stuff is all just like when I figure it out in my brain. So I wrote an article that's like how RLHF actually works, which is just the intuitions that I had put together in the summer about RLHF. And that was pretty well. And then I opportunistically wrote about Q star, which I hate that you have to do it, but it is pretty funny from a literature perspective and like opening high publishes on work that is very related to mathematical reasoning. So it's like, oh, you just poke a little around what they've already published. And it seems pretty reasonable, but we don't know. They probably just got like a moderate bump on one of their benchmarks. And then everyone lost their minds. It doesn't really matter.
**Swyx** (3:16)
Like this is why Sam Altman was fired. I don't know. Anyway, we're here to talk about RLHF 101 You did a presentation and I think you expressed some desire to rerecord it. And that's why I reached out on Twitter saying like, why not rerecord it with us? And then we can ask questions and talk about it.
**Nathan Lambert** (3:29)
Yeah, sounds good. I try to do it every six or 12 months as my estimated cadence, just to refine the ways that I say things and people will see that we don't know that much more, but we have a little bit of better way of saying what we don't know.
89 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000641340012