**Dwarkesh Patel** (0:00)
Today, I want to talk about what the compute situation for the labs will look like over the next few years. For the last three consecutive years, Anthropic's revenue has 10x year over year, and it's likely to do so again this year. So they ended last year with 9 billion in revenue. I think they'll probably end this year with somewhere between 100 billion to 150 billion dollars in revenue. Now, for this trend to continue, Anthropic would need to make one trillion dollars in revenue by the end of next year. Of course, there's no deep reason why this has to be true. It's a very wild conclusion, and it's ultimately a question of AI capabilities. Does AI get that useful by the end of next year? But suppose the trend does continue. Well, I want to think through what happens in that world. Now, the other big trend in AI is that lab compute only three X's year over year. For a lab to keep 10X in revenue year over year, while compute only three X's, one of the following three things needs to happen or some combination of the three needs to happen. One, lab margins have to increase. Two, the price of compute has to increase. Or three, the percentage of compute that lab spent on inference rather than training has to increase. My understanding is that basically all three of these things are already happening. With regards to the margins, Anthropics inference margins reportedly went from 40 percent to the middle of last year to upwards of 80 percent now available. With regards to compute, the spot prices for compute are more than 40 percent higher than they were in the February trough that we had earlier this year. And with regards to the share of compute that goes to training versus inference, in 2024, according to EPOC, OpenAI was spending just a quarter of its compute on inference and that number is likely closer to 50 percent, if not higher now. Now, Labs have deferred not to do this final thing, of increasing the share of compute they spend on inference.
The way the Labs see the world, the whole point of inference revenue is to help convince investors to give you more money in order to train the next bigger, better model. And if you're spending most of your compute on inference, you're basically declaring that AI progress has all and you're just now in the business of being a cloud provider. Now, this is a less compelling business than building AGI and so the Labs do not want to be in this business, nor do they think they're in this world. They think that within a year, they'll have built models that make the current ones look extremely shitty. But they need to invest a lot of their compute, the majority of their compute, into doing the training and experiments that are necessary to build the next model. So that leaves only two options for how you can get out of this gap between the fact that lab compute only increases 3x year over year, but revenue increases 10x.
Either the labs margins have to increase so that they get this surplus, or the price of compute has to increase so that everybody in the stack below the lab gets a surplus. It's not clear to me which world we end up in. Do we end up in a world where we go from 80% margins for some of the top models to greater than 90% margins if the lab margin effect dominates? Well, that would require the leading model to be so far ahead of the competition, because the nature of margins, why they exist in a market economy, is that the thing you are serving is so much better than what somebody else could go get and replace you on the market. But it's just really wild for me to consider that the margins for something like intelligence will be greater than 90%, and they don't get competitive at that level. So, that leaves only one other possibility of this escape valve between these two trends, which is that the price of compute has to increase. As I mentioned, this is already starting to happen, and the effect is even stronger when you look at the tranche of compute that the frontier labs actually need to accumulate, because they can't just go out and buy a spot instance. They need to make sure that they get enough scale to get really good efficiency and flexibility, and also that they have the kind of compute that lends itself to the security they need for their own weights and for their customer's information. So, I think a relevant case study here is to look at the computes that Google and Anthropic are renting from SpaceX. Google, for example, is paying $900 million a month for 110,000 GPUs that are a blend of GB 200s and GB 300s. The price that Google is paying here is 2X the spot price per hour for those GPUs. And that spot price itself is more than 40 percent higher than it would have been in February. I want to emphasize a key conclusion here, that as AI models get smarter, they will be better able to monetize the same amount of compute. If a true human level software engineer could run on an H100 equivalent, then at today's prices for software engineers, that H100 should rent for over 250K a year.
7 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID