**Dwarkesh** (0:00)
Today, we are interviewing Satya Nadella, we being me and Dylan Patel, who is founder of SemiAnalysis. Satya, welcome. Thank you.
**Satya Nadella** (0:08)
It's great. Thanks for coming over to Atlanta.
**Dwarkesh** (0:10)
Yeah. Thank you for giving us a tour of the new facility. It's been really cool to see. Absolutely. Satya and Scott Guthrie, Microsoft's EVP of Cloud and AI, give us a tour of their brand new Fairwater 2 data center, the current most powerful in the world.
**Scott Guthrie** (0:25)
We try to 10x the training capacity every 18 to 24 months.
So this would be effectively a 10x increase. 10x from what GPD 5 was trained with. So to put it in perspective, the number of the network optics in this building is almost as much as all of Azure across all our data centers two and a half years ago.
**Satya Nadella** (0:44)
It's five million network connections.
**Dwarkesh** (0:47)
You've got all this bandwidth between different sites in a region and between the two regions. So is this a big bet on scaling in the future that you anticipate in the future, there's going to be some huge model that needs to require two whole different regions, to train.
**Satya Nadella** (0:59)
The goal is to be able to kind of aggregate these flops for a large training job, and then put these things together across sites.
**Dwarkesh** (1:08)
Right.
**Satya Nadella** (1:08)
The reality is you'll use it for training, and then you'll use it for data gen, you'll use it for inference in all sort of ways. It's not like it's going to be used only for one workload forever.
**Scott Guthrie** (1:20)
Fairwater 4, which you're going to see under construction nearby, will also be on that one petabits network, so that we can actually link the two at a very high rate. Then basically, we do the AI WAN connecting to Milwaukee where we have multiple other Fairwaters being built.
**Satya Nadella** (1:35)
Literally, you can see the model parallelism and the data parallelism. It's built for essentially the training jobs, the pods, the super pods across this campus. Then with the WAN, you can go to the Wisconsin data center, and you literally run a training job with all of them getting aggregated.
**Scott Guthrie** (2:00)
What we're seeing right here is this is a cell with no servers in it yet, no racks.
**Dylan Patel** (2:04)
How many racks are in a cell?
**Scott Guthrie** (2:06)
We think about it. We don't necessarily share that per se, but we- That's the reason I asked. You'll see upstairs.
**Dylan Patel** (2:14)
I'll start counting.
**Scott Guthrie** (2:15)
You can start counting. We'll let you start counting.
**Dylan Patel** (2:16)
How many cells are there in this building?
**Scott Guthrie** (2:17)
That part also I can't tell you.
**Dylan Patel** (2:19)
Division is easy, right?
**Satya Nadella** (2:24)
My God, it's kind of loud.
**Dwarkesh** (2:27)
Are you looking at this like, now I see where my money is going.
**Satya Nadella** (2:30)
It's kind of like, I run a software company. Welcome to the software company.
**Dwarkesh** (2:35)
How big is the design space once we've decided to use the GB200s and NVLink? How many other decisions are there to be made?
**Satya Nadella** (2:41)
It is coupling from the model architecture to what is the physical plan that's optimized. And it's also scary in that sense, which is, hey, there's going to be a new chip that will come out, which obviously, I mean, you take Vera-Rubin Ultra. I mean, that's going to have power density that's going to be so different, but with cooling requirements that are going to be so different, right? So you kind of don't want to just build all to one spec. So that goes back a little bit to, I think the dialogue we'll have, which is you want to be scaling in time, as opposed to scale once, and then be stuck with it.
**Dylan Patel** (3:21)
When you look at all the past technological transitions, whether it be railroads or the internet or replaceable parts in drush utilization, the cloud, all of these things, each revolution has gotten much faster in the time it goes from technology discovered to ramp and pervasiveness through the economy. Many folks who have been on Dwarkesh's podcast believe this is sort of the final technological revolution in our transition, and this time is very, very different. At least so far in the markets, it's sort of in three years, we've already skyrocketed to hyperscalers are doing $500 billion of CapEx next year, which is a scale that's unmatched to prior revolutions in terms of speed. The end state seems to be quite different. Your framing of this seems quite different than sort of the I would say the AI bro, who is quite AGI is coming, and I'd like to understand that more.
77 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000736461892