The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman artwork

The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

The MAD Podcast with Matt Turck

July 23, 2026

AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act.
Speakers: Andrew Feldman, Matt Turck
**Andrew Feldman** (0:00)
This is the largest chip built in the history of the computer industry. It's 58 times larger than the GPU. And for AI, bigger chips process information more quickly, and therefore you get answers in less time. For AI world, big chips are undoubtedly the best way to go. There's no mode in inference. It takes you eight keystrokes to move from a GPU to us in the cloud. We solved a problem that nobody in the history of computing solved, and we delivered it in 2020, and nobody cared.
Nobody cared. Nobody bought it, and nobody cared. Everybody said we were crazy, it would never work, so then we built the next one.

**Matt Turck** (0:40)
Hi, I'm Matt Turck, welcome to the MAD Podcast. My guest today is Andrew Feldman, co-founder and CEO of Cerebras, the company that built the largest chip in the history of computing, and just pulled off the biggest semiconductor IPO of all time. Andrew has been everywhere talking about the headlines, the 20 billion dollar plus OpenAI deal, the IPO, but this conversation is a bit different. We started from what is a wafer and built up step by step, why GPUs struggle with fast inference, the three shortages nobody talks about, the decade in the desert when nobody wanted this chip, and why Andrew believes that CUDA is no longer a mode for Nvidia. If you want to actually understand the current chip landscape and how AI inference works at the silicon level, this episode is for you. Please enjoy this fantastic conversation with Andrew Feldman.

**Matt Turck** (1:31)
I thought a fun place to start would be to talk about speed. So has speed become the dominant conversation for AI today?

**Andrew Feldman** (1:41)
What happened, I think, was for a long time AI was sort of a novelty, right? It was like a parlor trick. It was cool, but not useful.
And what happened somewhere around the middle of 2025 was the AI got smart enough such that people began to use it. And we remember we make AI with training, but we use it with inference. And suddenly people wanted to use it. And the minute you want to use it, the minute it's productive, right, speed matters, right? Fast tokens are more productive. And so the conversation moved from everything else to, how do we make our inference faster? How do we deliver tokens more quickly? Because those are more productive tokens, we get more done in less time, and therefore it's more valuable.

**Matt Turck** (2:32)
Right. And what does speed mean?
Is that a question of token speed? Is that completion of the task? What's the right metric?

**Andrew Feldman** (2:39)
Sure. The right metric is tokens per second per user. That's how fast you get the first token all the way through the last token in your response. And it's true for queries from the chat, but it's also true for agentics loads. Right?
If there are sort of multi-cycle turns, waiting is amplified. And so what you want is blisteringly fast responses so that the AI feels like it's in real time. You can engage with this.

**Matt Turck** (3:16)
So it's the broadband moment for AI?

**Andrew Feldman** (3:20)
I think that's right. And I think that's a very good analogy. I think if you think of something like Netflix, when the Internet was slow, Netflix delivered DVDs and envelopes, you would get a DVD in an envelope. And when the Internet became fast, they didn't get more efficient at delivering DVDs and envelopes, they became a movie studio.
The speed enabled them to become something completely different. And that's what speed does in general and in particular for AI. It opens up a whole new domain, it allows you to use the AI differently, you will stay longer, you will come more often, and you work on harder problems.

**Matt Turck** (4:05)
Yep. So it's literally a question of the UX, right? That's just nobody wants to wait a few seconds.

**Andrew Feldman** (4:11)
That's right. I mean, how big is the market for slow search? How big is the market for dial-up is zero?
How big is... How long will you wait for a website to resolve? Will you wait eight seconds? Nobody waits. And so it's the exact same with AI.

**Matt Turck** (4:26)
So no more people waiting with their laptops open while the agent...

**Andrew Feldman** (4:30)
That's right. Well, it's running and running and running. I think that is not what people want.

**Matt Turck** (4:36)
Okay. Great. Wonderful. So I'd love to talk about the landscape of the chip industry right now to help people visualize and guard where you guys are. So there used to be basically this concept of one chip to do it all.

49 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000778023479