**Sarah Guo** (0:02)
Hey, everyone. Welcome to No Priors. I'm Sarah Guo.
**Elad Gil** (0:08)
I'm Elad Gil.
**Sarah Guo** (0:09)
This week on No Priors, we're back with another episode where we answer your questions about tech, AI, and everything in between.
**Elad Gil** (0:15)
I think we have a lot of different questions that people have brought up this week that they were hoping we could cover and some topics that we thought would be kind of interesting.
**Sarah Guo** (0:21)
I want to go to one of our listener questions, and I think a topic that's really popular with many of the companies that you and I work with in terms of access to computing for much smaller scale experiments.
**Elad Gil** (0:35)
What's going on with the GPU crunch?
**Sarah Guo** (0:37)
Yeah, the companies that you and I work with, many of them are companies that they need to use very specific infrastructure to train and serve large models. These work on GPUs, and the structure of the industry is like, it's just not very robust. You have a very small number of producers, NVIDIA and AMD generally, and then NVIDIA is very far ahead on the high-end processors that are most efficient for large-scale training and infrareds. Then you have the pandemic supply disruption, which we haven't fully recovered for. If you actually look at the supply chain, you go from the actual designers to the reliance on a few major foundries, like TSMC.
Expansion of this capacity is not easy. New fabs are billions of dollars. Yield is a very complicated thing. You can think of it as a massive precision manufacturing problem where temperature, pressure, chemical concentration, tool imperfections, new processes, materials issues, like anything can make production have lower yield or lower quality.
If you think about the speed with which the industry, driven by both large and small players, has decided that they want to do AI, the physical processes cannot keep up with that demand. It's as if half the companies in the world over a year-long period decided like, yeah, we need super computers, not super conductors, but gigantic networked GPUs.
**Elad Gil** (2:23)
What is the actual gap? To your point, it sounds like much of the AI world is dependent on GPUs in order to train and then do inference on these big AI models. The big suppliers are basically NVIDIA, AMD, and then there's a long tail of smaller folks.
What is the delta between the amount of capacity that exists today and that's needed? Are we off by 2X, 10X, some other number?
**Sarah Guo** (2:47)
It's hard to say because right now, there's no way to explore the price elasticity of these things, right? Just very specifically, the industry is kind of looking at deliveries in small quantity in September, larger quantities in December, January. Most of the large cloud providers are sold out for any scale for at least through April of next year.
And so, you have really interesting dynamics, like large cloud players who are the biggest consumers of these GPUs already, like a Microsoft going and buying from other providers for near-term supply, right? So, I think one question that I ask you is like, hey, do you think this is a long-term thing? Do you think it's a very short-term thing?
But I think it just goes back to the fundamental dynamics are, do you expect the demand for these chips to continue increasing at a pace that out-creases the ability to scale a very physical, like, real-world process, right? Just to even be more specific, one of the challenges, like I was talking to Jensen about this, and a bonder, like not part of the GPU itself, but like a critical tool in the manufacturing and assembly of GPUs is very specialized, So, I think we're seeing a lot of new things that are being specialized. And so, the ability to build any of these tools as well to enable these processes is a blocker. If you look at the demand from large labs today to continue increasing model scale and training time by magnitudes, I think it's hard to see that dynamic going away. What do you think?
**Elad Gil** (4:30)
I feel like there's a couple of different sort of second-order implications of the fact that we're seeing this giant GPU bottleneck.
I think the first one is that we're seeing new sort of models that are dependent on GPU access or ownership as ways to create all sorts of really interesting monetization and potentially eventually cloud services. So that's things like CoreWeave or FoundryML or other companies that are basically providing now GPUs in different ways, in some cases through aggregation or federating different sources of GPUs. In some cases, it's just having these large GPU clouds and being able to use them in really interesting ways.
18 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000624030307