Llama 2: The New Open LLM SOTA (ft. Nathan Lambert, Matt Bornstein, Anton Troynikov, Russell Kaplan, Whole Mars Catalog et al.) artwork

Llama 2: The New Open LLM SOTA (ft. Nathan Lambert, Matt Bornstein, Anton Troynikov, Russell Kaplan, Whole Mars Catalog et al.)

Latent Space: The AI Engineer Podcast

July 19, 2023

As first discussed on our May Emergency pod and leaked 4 days ago, Llama (renamed from LLaMA) was upgraded to Llama 2 (pretraining on 2 trillion tokens with 2x the context length - bigger than any dataset discussed in Datasets 101, and adding ~$20m of RLHF/preference annotation) and released for...
Speakers: Swyx, Alessio Fanelli, Simon Willison, Nathan Lambert, Matt Bornstein, Alex Volkov, Anton Troynikov, Russell Kaplan
**Swyx** (0:00)
Hello, hello, this is Swix, and I hope you like that new intro music that we have for the Emergency Pod.
For those of you who are new, we do Emergency Pods whenever there are big enough breaking news in the AI landscape, because we try to be the first place that all you AI engineers hear about news that might affect your day-to-day work. So a couple months ago in our No-Motes Emergency Pod that we did around the Google No-Motes demo, we actually talked a little bit about the rumors that Zuck would be considering releasing commercially a version of Llama. Four days ago, it was rumored and leaked in the press. Today at 9 AM, they released it.
You're probably listening to this on the Wednesday, so a day later. So usual MO about this is that we try to gather some guests, and then we try to talk through day one reactions from AI engineers. Today was a little bit special because we got some really special guests.
We had Nathan Lambert from Huggingface. He works at Huggingface as a machine learning researcher, and Huggingface were launch partners of MEDA's Llama2, which meant that they had early access. And so Nathan actually just dropped his in-depth paper review and summary, and he spent the most time with it. So we figured we should ask him the most number of questions because the rest of us were just reacting live to it. We also, which is a first for us, worked with Matt Bornstein of A16Z, which are also surprisingly launch partners of Llama2. They put up the first templates and the first playgrounds, Llama2.ai, which is super helpful for people trying this out for the first time. I also want to give a special shout out to my friend, Rogiko Radurovich. He couldn't join us today, but he spent a lot of time prepping the examples and talking with me through how to test Llama2 to compare its quality to Gpt3.5. We also had Anton, the CTO of Chroma, join to talk about the impact of Llama and open source retrieval augmented generation. And then finally, Russell Kaplan from Scale AI on how to fine tune Llama2. So a very guest-packed episode.
As always, it's a little bit awkward trying to be a moderator and participant in a Twitter space because you're always trying to see who goes first. But we've done our best to clean up the audio and make it an enjoyable listening experience or reading experience if you want to read on Substack. So let us know what you think. We tried to cover all the major issues and predictions with Llama2 and enjoy.

**Alessio Fanelli** (2:28)
There's not a single dull day in this space. I think when we started the podcast in January, a lot of people asked us, how long can you really do this? Just focusing on AI research and models. And I think the answer is clear now, a long time. So excited for this and excited to have Simon again. You're basically an honorary guest host of all of our Twitter spaces.

**Simon Willison** (2:49)
Cool, thank you. Now it's great to be here again.

**Alessio Fanelli** (2:52)
And Nathan, thanks for joining us. Actually share your write-up on Llama2 technical details with Sean this morning.
So it's great to have you here to dive into some of the details.

**Nathan Lambert** (3:02)
Yeah, it sounds good.
It's probably clear how Huggingface was trying to collaborate on releasing the model on the platform. So we ended up getting some early details, which made it a lot easier for me to cram study before the chaos hit.

**Alessio Fanelli** (3:17)
No, that's great. It's kind of what happened with the code interpreter episodes when Sean and I had access for about five hours and Simon was like, I've been playing with this for weeks and I had all the inside scoops. So I think this will be a good episode.
Maybe Nathan, you just want to give people a little bit of background on what you do at Huggingface and yeah, your experience with the Llama2 kind of preview.

**Nathan Lambert** (3:40)
Yeah, so I've been a researcher and helping lead reinforcement learning from human feedback efforts at Huggingface, which really means I do some research and I try to figure out how to fine tune models to do what people want. Generally, we're trying to operate in the scale a little bit smaller than what Meta is doing because we obviously don't have that kind of resources at a startup. So I do a lot of technical research and also try to actually engage and communicate that with the community.
And specifically to Llama, I think I was most interested on kind of the research side. I think the paper is a phenomenal artifact and it's clear that the model is really strong in a lot of areas and then kind of the big picture trends of where open source is going. Like this is a clear step in a direction that a lot of people wanted, but weren't sure if it was gonna happen.

73 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000621643513