**Dario Amodei** (0:00)
A generally well-educated human, that could happen in two or three years.
**Dwarkesh Patel** (0:05)
What does that imply for Anthropic when in two to three years, these Leviathans are doing like $10 billion training runs?
**Dario Amodei** (0:11)
The models, they just want to learn. And it was a bit like a Zen Cohen. I listened to this and I became enlightened.
The compute doesn't flow, like the spice doesn't flow.
It's like, you can't, like the blob has to be unencumbered, right? The big acceleration that happened late last year and the beginning of this year, we didn't cause that. And honestly, I think if you look at the reaction of Google, that that might be 10 times more important than anything else. There was a running joke. The way building AGI would look like is, you know, there would be a data center next to a nuclear power plant next to a bunker.
**Dwarkesh Patel** (0:44)
But now it's 2030 What happens next? What are we doing with a superhuman god?
Okay, today I have the pleasure of speaking with Dario Amodei, who is the CEO of Anthropic. And I'm really excited about this one. Dario, thank you so much for coming on the podcast.
**Dario Amodei** (0:59)
Thanks for having me.
**Dwarkesh Patel** (1:00)
First question, you have been one of the very few people who has seen scaling coming for years, more than five years. I don't know how long it's been, but ask somebody who's seen it coming, what is fundamentally the explanation for why scaling works? Why is the universe organized such that if you throw big blobs of compute at a wide enough distribution of data, the thing becomes intelligent?
**Dario Amodei** (1:21)
I think the truth is that we still don't know. I think it's almost entirely an empirical fact. I think it's a fact that you could kind of sense from the data and from a bunch of different places, but I think we don't still have a satisfying explanation for it. If I were to try to make one, but I'm just, I don't know, I'm just kind of waving my hands when I say this.
There's these ideas in physics around like long tail or power law of like correlations or effects.
And so like when a bunch of stuff happens, right? When you have a bunch of like features, you get a lot of the data in like kind of the early, you know, the fat part of the distribution before the tails. You know, for language, this would be things like, oh, I figured out there are parts of speech and nouns follow verbs. And then there are these more and more and more and more subtle correlations.
And so it kind of makes sense why there would be this, you know, every log or order of magnitude that you add, you kind of capture more of the distribution. What's not clear at all is why is it scale so smoothly with parameters? Why does it scale so smoothly with the amount of data? You can think up some explanations of why it's linear, like the parameters are like a bucket, and so the data is like water, and so size of the bucket is proportional to size of the water. But like, why does it lead to all this very smooth scaling? I think we still don't know. There's all these explanations. Our chief scientist, Jared Kaplan, did some stuff on like fractal manifold dimension that like you can use to explain it. So there's all kinds of ideas, but I feel like we just don't really know for sure.
**Dwarkesh Patel** (3:02)
And by the way, for the audience who is trying to follow along, by scaling, we're referring to the fact that you can very predictably see how if you go from GPT-3 to GPT-4, or in this case, Claude 1 to Claude 2, that the loss in terms of whether it can predict the next token scales very smoothly.
So, okay, we don't know why it's happening, but can you at least predict if empirically, here is the loss at which this ability will emerge, here is the place where this circuit will emerge. Is that at all predictable or are you just looking at the loss number?
**Dario Amodei** (3:30)
It is much less predictable. What's predictable is this statistical average, this loss, this entropy, and it's super predictable. It's like, you know, predictable to like sometimes even to several significant figures, which you don't see outside of physics, right? You don't expect to see it in this messy empirical field.
But actually specific abilities are very hard to predict. So, you know, back when I was working on GPT-2 and GPT-3, like when does arithmetic come in place? When do models learn to code? Sometimes it's very abrupt.
114 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000623806335