**Dylan Patel** (0:00)
This is the biggest change in human history maybe ever. What's about to happen with AI? This is the biggest revolution, bigger than industrial revolution. Jensen is very paranoid about losing. If he just kept making his mainline chip, people crush him on cost and performance. Acquiring Grok is how you get those resources to make more solutions for different parts of the market to stay king. At the end of the day, this is an economic war. If the US and the West win in AI, China will not rise to be the global hegemony. But without AI, China definitely will rise. They're just gonna outrun America.
**Matt Turck** (0:29)
Hi, I'm Matt Turck. Welcome back to The MAD Podcast. Today I'm joined by the one person Wall Street and Silicon Valley turn to when they need to cut through the hardware hype. Dylan Patel of SemiAnalysis. We dove into many of the most important topics of today. Nvidia's massive move to acquire Grok, the truth about the CapEx bubble, whether the US power grid can actually handle the AI boom and the geopolitical chess match playing out between the US and China. But I have to warn you, this conversation went off the rails in the best possible way and we ended up going into all sorts of fun tangents like the strange phenomenon of Chinese romance dramas set inside semiconductor factories and what it's really like when three AI-famous roommates live together in SF. Please enjoy this fantastic conversation with Dylan. Hey Dylan, welcome.
**Dylan Patel** (1:15)
Hello, how are you?
**Matt Turck** (1:16)
I'm great. I'd love to start with Grok and Nvidia since it's still fresh. So not so long ago Nvidia was saying that one GPU could do it all and now they're doing this acquisition slash non-exclusive deal with Grok. What does that mean from your perspective?
**Dylan Patel** (1:32)
It's very clear we're not sure where AI models are headed in terms of over the next few years. What happens to the architecture. But the thing that I think everyone has sort of agreed on is models are pretty autoregressive, right? Next token generation is like the thing. But beyond that, attention mechanisms change how it works. Everything changes, right? Could change. And so what's interesting is the reason Nvidia won is because they just took like the widest surface area bet, and then people kept developing models on that, and that kind of shape worked. But now the workload is so large that there is room for specialization that will give you 10x increases in certain domains, right? In a general purpose workload, croc, croc, doesn't work, right? You know, it can't train, it can't, you know, it can't inference really, really large models. Cost efficiently, right? You can't serve many, many, many users. But what it can do is it can go screamingly fast, right? Same with the Cerebras OpenAI deal. But that's like one workload, right? Very decode focused, right? Doing autoregressive tokens in a single stream super fast. Another direction AI models could head, right? We don't know, are models going to think in one token stream? Or is it actually they're constantly context switching, right? And they're going from, they have this humongous, humongous context and they're generating in multiple parallel streams, right? And so Google and OpenAI have both released mechanisms of this with their pro models, where the model actually doesn't just have one single chain of thought for reasoning, it has multiple, right? And then I don't exactly, like, you know, and how they choose which one and what the final answer to you delivers is an area of research. But there is room for that kind of chip, right? Something that works on very parallel a lot of streams of chain of thought. And maybe the latency requirements are not as crazy, right? Maybe you don't want to go blindingly fast, right? Maybe you're okay with it being, you know, because I can spin up a hundred parallel, you know, streams of thought or agents or whatever you want to call them. Maybe I care a lot about cost there. And because it's a hundred in parallel instead of one going super, super fast, it's not as deep, right? The tree searcher, the depth of the inference is not as deep, but it is much wider. You know, there's other parts of inference. Hey, process creating the KV cache. So Nvidia has a chip for that, right? That's the CPX. So they've made the CPX, they bought Grok for decode, and then they still have their general purpose GPU. So they're kind of trying to cover their bases because unlike the first wave of AI chip companies where they sort of just made chips and then tried to figure out where it would work, right? They had a thesis, Grok and Cerebras both, as well as Samba Nova, right? Which has put a lot of memory on the chip and not necessarily in the case of Cerebras and Grok, no memory off chip. And in the case of Samba Nova, less memory off chip or slower memory off chip with higher capacity, you know, they sort of all made similar bets in that direction. And it didn't work for a while until it kind of did, right? Because there was a workload that now necessitates it. Nvidia recognizes they're the leader, they're at the tent pole. Hey, in one respect, they can just run faster than everyone, but it's kind of hard to be 2x better than Google or OpenAI or whoever else's internal chip, right? To justify their, you know, 75% plus margins, right? And then they have to be 2x to 4x better to justify, 4x better to justify their margins because that's what they're charging above cogs. You know, the question is, what architecture will deliver that? Well, yes, keep the programmability of their GPUs is great for training and for a lot of workloads. But, you know, guess what?
80 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000748369780