Mark Zuckerberg — Llama 3, $10B models, Caesar Augustus, & 1 GW datacenters artwork

Mark Zuckerberg — Llama 3, $10B models, Caesar Augustus, & 1 GW datacenters

Dwarkesh Podcast

April 18, 2024

Mark Zuckerberg on: - Llama 3 - open sourcing towards AGI - custom silicon, synthetic data, & energy constraints on scaling - Caesar Augustus, intelligence explosion, bioweapons, $10b models, & much more Enjoy! Watch on YouTube. Listen on Apple Podcasts, Spotify, or any other podcast platform.
Speakers: Dwarkesh Patel, Mark Zuckerberg
**Dwarkesh Patel** (0:00)
Mark, welcome to the podcast.

**Mark Zuckerberg** (0:01)
Hey, thanks for having me. Big fan of your podcast.

**Dwarkesh Patel** (0:03)
Oh, thank you. That's very nice of you to say.
Okay, so let's start by talking about the releases that will go out when this interview goes out. Tell me about the models, tell me about MetaAI. What's new, what's exciting about them?

**Mark Zuckerberg** (0:15)
Yeah, sure. So I think the main thing that most people in the world are going to see is the new version of MetaAI. So the most important thing about what we're doing is the upgrade to the model. We're rolling out Llama 3, we're doing it both as open source for the dev community, and it is now going to be powering MetaAI. So there's a lot that I'm sure we'll go into around Llama 3, but the bottom line on this is that with Llama 3, we now think that MetaAI is the most intelligent AI assistant that people can use that's freely available. We're also integrating Google and Bing for real-time knowledge. We're going to make it a lot more prominent across our apps. So basically, at the top of WhatsApp and Instagram and Facebook and Messenger, you'll just be able to use the search box right there to ask any question. There's a bunch of new creation features that we added that I think are pretty cool, that I think people enjoy. I think animations is a good one. You can basically just take any image and animate it. But I think one that people are going to find pretty wild is, it now generates high-quality images so quickly. I don't know if you've gotten a chance to play with this, that it actually generates it as you're typing and updates it in real-time. So you're typing your query and it's honing in on, and it's like, okay, here, show me a picture of a cow in a field with mountains in the background. It's just like everything's popular. Eating macadamia nuts, drinking beer, and it's updating the image in real-time. It's pretty wild. I think people are going to enjoy that.
That I think is what most people are going to see in the world. We're rolling that out, not everywhere, but we're starting in a handful of countries, and we'll do more over the coming weeks and months.
That I think is going to be a pretty big deal. I'm really excited to get that in people's hands. It's a big step forward for Met AI.
But I think if you want to get under the hood a bit, the Llama 3 stuff is obviously the most technically interesting.
For the first version, we're training three versions, an 8 billion and a 70 billion, which we're releasing today, and a 405 billion dense model, which is still training. So we're not releasing that today. But the 8 and 70, I'm pretty excited about how they turned out. I mean, they're leading for their scale. We'll release a blog post with all the benchmarks so people can check it out themselves. And obviously, it's open source, so people get a chance to play with it. We have a roadmap of new releases coming that are going to bring multimodality, more multilinguality, bigger context windows to those as well.
And then hopefully, sometime later in the year, we'll get to roll out the four or five, which I think is in training. It's still training. But for where it is right now in training, it is already at around 85 MMLU.
And we expect that it's going to have leading benchmarks on a bunch of the benchmarks. So I'm pretty excited about all of that. I mean, the 70 billion is great, too. I mean, we're releasing that today. It's around 82 MMLU and has leading scores on math and reasoning. So I mean, I think just getting this in people's hands is going to be pretty wild.

**Dwarkesh Patel** (3:42)
Oh, interesting. Yeah, that's the first time hearing it in a benchmark. That's super impressive.

**Mark Zuckerberg** (3:45)
Yeah, the 8 billion is nearly as powerful as the biggest version of Llama 2 that we released. So it's like the smallest Llama 3 is basically as powerful as the biggest Llama 2

**Dwarkesh Patel** (3:59)
OK, so before we dig into these models, I actually want to go back in time.
2022 is, I'm assuming, when you started acquiring these H100s.
Or you can tell me when, where you're like stock price is getting hammered. People are like, what's happening with all this capex? People aren't buying the metaverse. And presumably you're spending that capex to get these H100s. How back then, how did you know to get the H100s? How did you know we'll need the GPUs?

68 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000652877239