The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis artwork

The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis

Latent Space: The AI Engineer Podcast

November 17, 2023

This episode came together at ~4 hrs notice since Dylan had just landed in SF and we had to setup quickly; you might notice some small audio issues in some segments, we apologize. We’re currently building our own podcast studio for 2024!
Speakers: Alessio, Swyx, Dylan Patel
**Alessio** (0:07)
Hey, everyone, welcome to the Latent Space Podcast. This is Alessio, Partner and CTO of Residence and Decibel Partners. I'm joined by my co-host, Swix, founder of Small AI.

**Swyx** (0:16)
And today we have Dylan Patel, and welcome. So you are the author of the extremely popular SemiAnalysis blog. We have both had a little bit of claim to fame in breaking details of GPT-4. George Hotz came on our pod and talked about the mixture of experts thing, and then you had a lot more detail on it.

**Dylan Patel** (0:29)
Well, to be clear, I talked about mixture of experts in January, it's just people didn't really notice it. I guess, I don't know.

**Swyx** (0:35)
You went into a lot more detail, and I'd love to dig into some of that. Yeah.

**Dylan Patel** (0:38)
Thank you so much. I've been doing consulting in the industry, it's semiconductor industry since 17 2021 got bored, and in November, I started writing a blog, and then like 2022 was good, and I started hiring folks from my firm, and then all of a sudden, 2023 happens, and it's like the perfect intersection. I used to do data science, but not like AI, not really, like multivariable progression is not AI, right? But also I've been involved in the semiconductor industry for a long, long time, posting about it online since I was 12, right? And all of a sudden, this all kind of came to fruition. So it's cool to have the blog sort of blow up in that way.

**Swyx** (1:10)
I used to cover semis at Baleazni as well. And it was for a long time, it was just the mobile cycle.
And then a little bit of PCs, but like not that much. And then maybe some cloud stuff, like public cloud semiconductor stuff, but it really wasn't anything until this wave. And I was actually listening to you on one of the previous podcasts that you've done. And it was surprising that high-performance computing also kind of didn't really take off. Like AI is just the first form of high-performance computing that worked.

**Dylan Patel** (1:37)
One of the theses I've had for a long time that I think people haven't really caught on, but it's coming to fruition now, is that the largest tech companies in the world, their software is important, but actually having and operating a very efficient infrastructure is incredibly important. And so, you know, people talk about, you know, hey, Amazon is great, AWS is great, because yes, it is easy to use, and they've built all these things. But behind the scenes, they've done a lot on the infrastructure that is super custom, that Microsoft, Azure, and Google Cloud just don't even match in terms of efficiency. If you think about the cost to rent out SSD space, so the cost to rent, you know, offer a database service on top of that, obviously, a cost to rent out a certain level of CPU performance, Amazon has a massive advantage there. And likewise, like Google spent all this time doing that in AI, right? With their TPUs and infrastructure there and optical switches and all this sort of stuff. And so in the past, it wasn't immediately obvious. I think with AI, especially like, how scaling laws are going, it's like incredibly important for infrastructure is like so much more important. And then like, when you just think about software cost, right, like the cost structure of it, there was always a bigger component of R&D and like SaaS businesses, you know, all over SF, all these SaaS businesses did crazy good because they just start as they grow. And then all of a sudden they're so freaking profitable for each incremental new customer. And AI software looks like it's going to be very different in my opinion, right? Like the R&D cost is much lower in terms of people, but the cost of goods sold in terms of actually operating the service, I think will be much higher. And so in that same sense, infrastructure matters a ton.

**Swyx** (3:02)
And I think you wrote once that training costs effectively don't matter.

**Dylan Patel** (3:06)
Yeah, in my opinion, I think that's a little bit spicy, but yeah, it's like training costs are irrelevant, right? Like GPT-4, right? Like 20,000 A100s, that's like, I know it sounds like a lot of money, it's child's play.

**Swyx** (3:17)
500 million all in, is that a reasonable estimate?

**Dylan Patel** (3:18)
The supercomputer, it's slightly more, but yeah, I think the 500 million is a fair enough number. I mean, if you think about just the pre-training, right? Three months, 20,000 A100s at a dollar an hour is like, that is way less than 500 million, right? But of course there's data and all this sort of stuff.

54 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000635213033