**Brad Gerstner** (0:00)
Elon is building a much, much bigger cluster to train a much, much bigger model, as is OpenAI, as is Zuckerberg. I mean...
**Bill Gurley** (0:08)
Well, Sam just said the bigger the models aren't the problem.
**Brad Gerstner** (0:11)
Well, I mean, he may be doing the same game that everybody else is doing, Bill, and trying to throw everybody off the scent.
Great to have you guys, what a week. It's been nuts, there's so much to talk about. And we have our good buddy, Sunny.
**Bill Gurley** (0:37)
How you doing, Sunny?
**Brad Gerstner** (0:38)
Sundeep Madra in the house. Sunny, somebody Bill and I go to often when we're talking through all things AI. Currently at Groq, working on the inference cloud.
So you're deep in thinking about all these things, all these models, AI. And we're gonna talk a lot about models today with the release of Llama 3 So it's good to have you, Sunny.
**Sunny Madra** (0:57)
Good to be here, thanks guys.
**Brad Gerstner** (0:58)
Bill, why are you in town?
**Bill Gurley** (1:00)
Board meetings, a couple of board meetings.
**Brad Gerstner** (1:02)
Good to have you. I always like doing this in person.
**Bill Gurley** (1:04)
Yeah, I know you do.
**Brad Gerstner** (1:04)
Okay, so let's roll.
**Sunny Madra** (1:06)
They're my favorite episodes when you do them live.
**Brad Gerstner** (1:08)
There we go, there we go.
So models, models, models, models. If AI is the next big thing, then this felt like another really important week. I mean, we got models being dropped by Meta with Llama 3 That was the one that was really, the category five earthquake, Microsoft, Snowflake, everybody seems to be out with a new model. But let's start with Zuck.
Huge Llama 3 unveiling three distinct models, an 8 billion, a 70 billion, and a 405 billion parameter model, which is still training and still learning, they're telling us, which is pretty fascinating. But what seems to have shocked the market is that Meta could pack so much intelligence into such a small model. And so both models quickly shot up the rankings this week.
We have a screenshot here of that. Of course, the 405 is still training, and there have been some hints out of a recent podcast with Zuck and Dworkesh about it may, in fact, kind of come in at the top of the polls. We'll see. It's probably going to train for another couple months. But I'd love to hear from both of you guys. What were your big takeaways from the launch of Llama3? And maybe start with you, Sunny. Walk us through kind of just the what and the how of Llama3 and why it really kind of shook things up.
**Sunny Madra** (2:28)
Yeah, I would say, you know, the biggest impact of Llama3 is its capabilities and at the size. And what, you know, Zuck shared in that interview was that they basically took the model and kept training it past the chinchilla point.
And so really by doing that, which is generally considered like sort of the point of diminishing returns, they were able to pack much more information and much more capability into this model with the same data set.
**Brad Gerstner** (2:56)
So just for everybody listening, so the chinchilla point, if I understand it correctly, right, it's the byproduct of this paper out of Google, which basically talked about the relationship between the optimal amount of data to use for a certain amount of compute.
But in the case of Meta, when they were training Lama 3, they were basically continued with these forward passes of the data. So they were curating the data, refining the data, pushing it back into the model. And I think several people who are working on pre-training at Meta said they were even surprised that it was still learning when they took it offline on that data.
**Sunny Madra** (3:35)
Yeah, and they only took it offline to reallocate the resources to 405 and other efforts.
**Brad Gerstner** (3:41)
And I think he said Lama 4
**Sunny Madra** (3:42)
And Lama 4
**Brad Gerstner** (3:43)
So the rate of innovation is certainly not slowing down there. So a 15 trillion parameter model.
**Sunny Madra** (3:48)
15 trillion tokens used to train it.
**Brad Gerstner** (3:51)
15 trillion tokens used to train, you know, the model.
I know at Groq, you guys are deploying Lama 3 I think you deployed it the same day that it came out. So how important is this? How important a development is it in the world of models?
**Sunny Madra** (4:09)
Well, really, you know, Zuck came out and threw down for the entire world folks that are building models.
55 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000653976211