**Dwarkesh Patel** (0:00)
All right, Mark, thanks for coming on the podcast again.
**Mark Zuckerberg** (0:02)
Yeah, happy to do it. Good to see you.
**Dwarkesh Patel** (0:04)
You too. Last time you were here, you had launched Llama 3 Yeah. Now you've launched Llama 4
**Mark Zuckerberg** (0:09)
Well, the first version.
**Dwarkesh Patel** (0:10)
That's right. What's new? What's exciting? What's changed?
**Mark Zuckerberg** (0:12)
Oh, well, I mean, the whole field is so dynamic. So I mean, I feel like a ton has changed since the last time that we talked. MetaAI has almost a billion people using it now monthly. So that's pretty wild. And I think that this is going to be a really big year on all of this, because especially once you start getting the personalization loop going, which we're just starting to build in now, really, from both the context that all the algorithms have about what you're interested in feed, and all your profile information, all the social graph information, but also just what you're interacting with the AI about. I think that's just going to be kind of the next thing that's going to be super exciting. So really big on that. The modeling stuff continues to make really impressive advances too, as you know. The Llama 4 stuff, I'm pretty happy with the first set of releases.
We announced four models, and we released the first two, the Scout and Maverick ones, which are kind of like the mid-sized models, mid-sized to small. It's not like... Actually, the most popular Llama 3 model was the 8 billion parameter model. So we've got one of those coming in the Llama 4 series too. Our internal code name for it is Little Llama. But that's coming probably over the coming months. But the Scout and Maverick ones, they're good. They're some of the highest intelligence per cost that you can get of any model that's out there, natively multimodal, very efficient, run on one host. Designed to just be very efficient and low latency for a lot of the use cases that we're building for internally and that's our whole thing. We basically build what we want and then we open source it, so other people can use it too. So I'm excited about that. I'm also excited about the behemoth model, which is coming up. That's going to be our first model that is at the frontier. I mean, it's like more than two trillion parameters. So it is, I mean, it's, you know, as the name says, it's quite big. So we're kind of trying to figure out how we make that useful for people. It's so big that we've had to build a bunch of infrastructure just to be able to post-train it ourselves. And we're kind of trying to wrap our head around, how does the average developer out there, how are they going to be able to use something like this? And how do we make it so it can be useful for distilling into models that are of reasonable size to run? Because you're obviously not going to want to run, you know, something like that in a consumer model.
But yeah, I mean, there's a lot to go. I mean, as you saw with the Llama 3 stuff last year, the initial Llama 3 launch was exciting, and then we just kind of built on that over the year. 3.1 was when we released the 405 billion model, 3.2 is when we got all the multimodal stuff in. So we basically have a roadmap like that for this year too. So a lot going on.
**Dwarkesh Patel** (3:14)
I'm interested to hear more about it. There's this impression that the gap between the best closed source and the best open source models has increased over the last year, where I know the full family of Llama 4 models is not yet. But Llama 4 Maverick is 35 on Chatbot Arena. And on a bunch of major benchmarks, it seems like O4 Mini or Gemini 2.5 Flash are beating Maverick, which is in the same class. What do you make of that impression?
**Mark Zuckerberg** (3:41)
Yeah, well, OK, there's a few things. I actually think that this has been a very good year for open source overall. Right, if you go back to where we were last year, what we were doing with Llama was the only real super innovative open source model. Now you have a bunch of them in the field. And I think in general, the prediction that this would be the year where open source generally overtakes closed sources, the most used models out there, I think is generally on track to be true. I think the thing that's been an interesting surprise, I think, positive in some ways, negative in others, but I think overall good, is that it's not just Llama. There are a lot of good ones out there. So I think that that's quite good. Then there's the reasoning phenomenon, which you basically are alluding to with talking about 3 and 4 and some of the other models. I do think that there is this specialization that's happening where if you want a model that is sort of the best at math problems, or coding, or different things like that, I do think that these reasoning models with a lot of the ability to just consume more test time or inference time compute in order to provide more intelligence is a really compelling paradigm. But for a lot of the applications that, and we're going to do that too. We're building a Llama 4 reasoning model and that'll come out at some point.
64 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000705423020