**Josh** (0:00)
Nvidia just announced the key to one of AI's biggest problems, cost. Companies spend between tens to hundreds of millions of dollars every year and are running out of money. Google, for the first time yesterday, reported a negative cash flow on their quarterly earnings. They're officially losing more money than they are making. That's because they're spending so much on AI infrastructure. Nvidia's new Vera Rubin officially got stood up yesterday, and it saves you 10X on the tokens that you spend, which means you can have a 10X better model for the same cost. There's all of this and Google's new Gemini 3.6 models, which is very, very underwhelming on today's roundup.
**Ejaaz** (0:37)
Yeah, so we have to start with the big news of the day, which is that Vera Rubin is actually shipping. And for those who are not familiar, Jensen came on stage at Nvidia GTC and announced these a few months ago. The news today is that they're finally live and they're actually operational, and we started to get an idea of what these things look like. And I want to preface this section with the idea that we just covered an episode yesterday about how GPT-6 or whatever the new internal OpenAI model was. They broke out of internal air-gapped containment and hacked into a public-facing website into their production database to steal secrets. That was on Blackwell chips. The new chips are Vera Rubin. No models have been trained on Vera Rubin chips, but the math behind how much more powerful they are is so unbelievably impressive. It's like, oh my god, this feels like the end game. It's like, once these Vera Rubin chips come online at scale, what on earth are these models going to look like if we already have Fable and GPT-6 class models? It's going to be pretty wild for some numbers.
10x, the efficiency, which is crazy. So for every single megawatt you put into a GPU, it will give you 10 times the amount of tokens. This is a huge unlock for a lot of models.
The second thing is in terms of density of transistors, there's 336 billion, that is a 62% increase over the Blackwell GB300. And this is on TSMC's 3 nanometer technology, which is basically the cutting edge. There's 22 terabytes per second of memory that is going through this whole thing. And basically, they ran out of room for a single piece of silicon, and they glued these two maxed out dyes together and call it one. So that's kind of where we are, is like this chip is going to be 10x more performant per watt, and it's going to have a just unbelievable baseline relative to the GB300. A lot of people are saying about 4x the baseline. So imagine what we get when these models are trained on not only 4x the baseline, but also efficiency improvements in terms of algorithms. Like the next generation of software built on these things is going to be a monster.
**Josh** (2:40)
And I want to translate what this means for the wider market and for Nvidia stock, which has pretty much just been flat for the last like five months. I think this is Nvidia's star moment for this year. Now, the reason is it's really costly to train and inference models these days. And so any way that a company, an enterprise that is spending 10 to hundreds of millions of dollars every year can save money is a big deal. Now, a big way that they can save money is infrastructure. Right now, people are paying around 50 to 60% profit margin on every dollar spent on an AI token to Nvidia. They are like monopolies in this way. So Nvidia giving them a 10x efficiency increase means that you can create or run your Fable 5 model, your local open source model, Kimi K3, at one tenth of the cost purely because of infrastructure. Now, the way that this GPU works, you can't just run one Vera Rubin on its own. You need to run 72 of them in one server rack. And they're serviced by 32 of Nvidia's CPUs. Now, the reason why they have so many CPUs is because when you're running these AI models these days, you're not only running one instance of an AI model, you're running many. It's called AI agents, and they need to access different tools. That's what the CPUs are for. So it's this collective thing with the software and hardware integration, the way it's set up, that makes Nvidia's GPUs so so effective. And so if I had to make a call here, what is their most direct competitor, Josh? It has to be Google. Google has tried to go for the Nvidia throne with their own custom-made GPUs called TPUs, their Tensor Processing Units, and they had their quarterly earnings results yesterday, and they made a significant chunk of money on TPUs, but they also announced that they're going to be releasing a new chip, and I just don't think that they can catch up to Nvidia. So Nvidia is running away with it, and this is a huge advancement for AI models in general.
20 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000778203157