**Dylan Patel** (0:00)
Nvidia is going to have better networking than you, they're going to have better HBM, they're going to have better process node, they're going to come to market faster, they're going to be able to ramp faster, they're going to have better negotiations with whether it's TSMC or SK Hynix and the memory and Silicon side or all the rack people or like copper cables, everything, they're going to have better cost efficiency. So you can't just like do the same thing as Nvidia. You have to really leap forward in some other way. You have to be like 5X better.
**Erik Torenberg** (0:23)
The AI race isn't just about models. It's also about the infrastructure underneath them, chips, data centers, power, networking, and the economics that determine who can keep scaling. In this conversation, SemiAnalysis co-founder Dylan Patel joins Erin Price-Wright, Guido Appenzeller and me to discuss the state of AI hardware, why Nvidia remains so difficult to compete with, and how companies like Google, Amazon, Meta and OpenAI are approaching the next generation of AI infrastructure.
We also explore custom silicon, AI economics, robotics, export controls, and what founders and investors should be paying attention to as the compute race accelerates.
**Erin Price-Wright** (1:09)
Dylan, welcome to the podcast.
**Guido Appenzeller** (1:11)
Thank you for having me.
**Erin Price-Wright** (1:11)
We've been trying to get you for a while, you're a busy man, but it worked out. Guido, why don't you introduce why we're so excited to have Dylan on the podcast and what we're excited to discuss.
**Guido Appenzeller** (1:19)
I think Dylan, you've done an exceptional job in covering what's happening in the AI hardware space, AI semi-space, and now more and more data center space as well. Just looking at it, currently, the most valuable company on the planet is an AI semi-company.
The biggest IPO so far in AI was an AI cloud company. This is currently where it's happening. In any gold rush in the early days is the pigs and truffles that make money. I think this is the stage that we're in. So, I'm super excited to have you here today.
**Dylan Patel** (1:46)
Awesome.
**Guido Appenzeller** (1:47)
Thank you.
**Dylan Patel** (1:47)
Happy to talk about my favorite topics.
**Erin Price-Wright** (1:50)
Amazing. Well, maybe let's start with GPT-5. We just had some of the research for Christina and Isabella on here last week. You said it was disappointing. Why don't you share your reactions or what capabilities you were hoping to see?
**Dylan Patel** (2:00)
I think it depends on what tier of user you are. Right, if you're just using GPT-5 and before you were $20 or $200 a month subscriber, you no longer have access to 4.5, which in my opinion is still a better pre-trained model for certain things, or you no longer have access to 03, which would think for 30 seconds on average maybe, right? Whereas GPT-5, even when you're using thinking, only thinks for like 5 to 10 seconds on average, right? Which is an interesting sort of phenomenon, right? But basically like GPT-5 is not spending more compute per se. The model did get a little bit better on a vanilla basis, right? 4 to 5 is actually quite a bit better. But when you think about, you know, what is this curve of intelligence, right? It's like the more compute you spend, the better the model gets. And that's whether it's a bigger model, which GPT-5 isn't, right? You can see it's not a bigger model. It's roughly the same size, you know, or you think more, right? But again, like, this is something that OpenAI's first thinking models, you know, the first few generations of 0.1, 0.3 would think for a long time and waste a lot of tokens, if you will. And when you look at, for example, Anthropics thinking models, even when you put them in thinking mode, they think a lot less, right? To get to the same results or better results, right? As OpenAI was. And so OpenAI, I think, like, optimized a lot of, like, well, if I ask, like, I think the silliest one I had asked was like, I asked 0.3 once, is pork red meat or white meat? And it thought for like 48 seconds, I was like, what are you doing? Like, this should just, like, tell me the answer. And so, like, the nice thing is that GPT-5 will think a lot less, even if you select thinking manually, but more importantly, they have the sort of auto functionality, the router, which lets them decide whether or not, hey, do I route to the regular model? Do I route to maybe many if you're at a rate limits? Or do I route to thinking? Right? And how much do I think? But in general, the thinking model will think less.
64 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000776888410