NVIDIA's Jensen Huang on AI Chip Design, Scaling Data Centers, and his 10-Year Bets artwork

NVIDIA's Jensen Huang on AI Chip Design, Scaling Data Centers, and his 10-Year Bets

No Priors: Artificial Intelligence | Technology | Startups

November 7, 2024

In this week’s episode of No Priors, Sarah and Elad sit down with Jensen Huang, CEO of NVIDIA, for the second time to reflect on the company’s extraordinary growth over the past year. Jensen discusses AI’s takeover of datacenters and NVIDIA’s rapid development of x.AI’s supercluster.
Speakers: Elad Gil, Jensen Huang, Sarah
**Elad Gil** (0:05)
Hi, listeners, and welcome to No Priors. Today, we're here again, one year since our last discussion with the one and only Jensen Huang, founder and CEO of Nvidia. Today, Nvidia's market cap is over $3 trillion, and it's the one literally holding all the chips in the AI revolution. We're excited to hang out in Nvidia's headquarters and talk all things frontier models and data center scale computing and the bets Nvidia is taking on a 10-year basis. Welcome back, Jensen. 30 years in to Nvidia and looking 10 years out, what are the big bets you think are still to make? Is it all about scale up from here? Are we running into limitations in terms of how we can squeeze more compute memory out of the architectures we have? What are you focused on?

**Jensen Huang** (0:47)
Well, if we take a step back and think about what we've done, we went from coding to machine learning, from writing software tools to creating AIs, and all of that running on CPUs that was designed for human coding, to now running on GPUs designed for AI coding, basically, machine learning. So the world has changed. The way we do computing, the whole stack has changed, and as a result, the scale of the problems we could address has changed a lot. Because if you could paralyze your software on one GPU, you've set the foundations to paralyze across a whole cluster, or maybe across multiple clusters or multiple data centers. And so I think we've set ourselves up to be able to scale computing at a level and develop software at a level that nobody's ever imagined before. And so we're at the beginning of that.
Over the next 10 years, our hope is that we could double or triple performance every year at scale, not at chip, at scale. And to be able to therefore drive the cost down by a factor of two or three, drive the energy down by a factor of two or three every single year. When you do that every single year, when you double or triple every year, in just a few years it adds up. So it compounds really, really aggressively. And so I wouldn't be surprised if, you know, the way people think about Moore's Law, which is a 2X every couple of years, you know, we're going to be on some kind of a hyper Moore's Law curve. And I fully hope that we continue to do that.

**Sarah** (2:29)
What do you think is the driver of making that happen even faster than Moore's Law? Because I know Moore's Law was sort of self-reflexive, right? It was something that he said and then people kind of implemented it to make it happen.

**Jensen Huang** (2:39)
The two fundamental technical pillars, one of them was Dennard scaling and the other one was Carver Mead's VLSI scaling. And both of those techniques were rigorous techniques, but those techniques have really run out of steam. And so now we need a new way of doing scaling. You know, obviously the new way of doing scaling are all kinds of things associated with co-design. Unless you can modify or change the algorithm to reflect the architecture of the system, or change and then change the system to reflect the architecture of the new software and go back and forth. Unless you can control both sides of it, you have no hope. But if you can control both sides of it, you can do things like move from FP64 to FP32 to BF16 to FP8 to FP4 to who knows what, right? And so I think that co-design is a very big part of that. The second part of it, we call it full stack. The second part of it is data center scale. Unless you could treat the network as a compute fabric and push a lot of the work into the network, push a lot of the work into the fabric, and as a result, you're compressing, doing compressing at very large scales. And so that's the reason why we bought Melanox and started fusing InfiniBand and NVLink in such an aggressive way. And now look where NVLink is going to go. The compute fabric is going to scale out what appears to be one incredible processor called a GPU. Now, we'll get hundreds of GPUs that are going to be working together. Most of these computing challenges that we're dealing with now, one of the most exciting ones, of course, is inference time scaling. It has to do with essentially generating tokens at incredibly low latency. Because you're self-reflecting, as you just mentioned. I mean, you're going to be doing tree serves, you're going to be doing chain of thought, you're going to be doing probably some amount of simulation in your head, you're going to be reflecting on your own answers. Well, you're going to be prompting yourself and generating text to your silently and still respond, hopefully in a second. Well, the only way to do that is if your latency is extremely low. Meanwhile, the data center is still about producing high throughput tokens because you still want to keep cost down, you want to keep the throughput high, you want to generate a return. So these two fundamental things about a factory, low latency and high throughput, they're at odds with each other. In order for us to create something that is really great in both, we have to go invent something new and NVLink is really our way of doing that.

27 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000676047524