**Elad Gil** (0:05)
The shortage of compute for AI, also known as the GPU crunch, is increasingly impacting the AI companies big and small. It is delaying training runs and launches from multiple players in the generative AI world. One company coming to the rescue is Cerebras Systems, which is developing the largest computer chip and one of the fastest purpose-built AI processors ever. This week, Sarah and I are joined by Andrew Feldman, CEO of Cerebra Systems.
Andrew is one of the few entrepreneurial veterans in the semiconductor world. He previously started CMicro, a pioneer of energy-efficient, high-bandwidth microservers. CMicro was acquired by AMD in 2012
Andrew, thank you so much for joining us today.
**Andrew Feldman** (0:44)
Elad and Sarah, thank you so much for having me.
**Elad Gil** (0:47)
I think you all recently announced that Cerebra has closed a $100 million deal with G42 to develop one of the largest AI supercomputers in the world.
**Andrew Feldman** (0:55)
We did announce a strategic partnership with a group called G42, and we announced that we were building nine supercomputers. Each supercomputer would be four exaflops of AI compute, so in total, 36 exaflops of AI compute.
That was extraordinarily exciting. When you encounter a partner that shares your vision and wants to build with you, and you get to start building the biggest computer on Earth, it doesn't get better than that.
**Elad Gil** (1:27)
And I think in general, you all were very forward thinking and early to identifying AI as a really important market for custom semiconductors. Could you tell us a little bit more about your thinking early on as you started Cerebros and why you focused on AI many years ago?
**Andrew Feldman** (1:41)
In late 2015, five of us started meeting regularly. I think we were meeting in Sarah's offices, actually, at that point. All of us had worked together at our previous company, and we began working on ideas.
And one day, our CTO, Gary, leaned back and said, why would a machine built for pushing pixels to a monitor be ideal for AI? Then he said, wouldn't it be serendipitous if 20 years of optimizing a part for one job left it really well suited for another? That got us excited. We began looking at GPUs, looking at AI work, and by early 2016, we decided that we could build a better part for this work. Our strategy was not to build a little bit better, but it was to try and do something vastly better. We went to a technology that had never worked before called wafer scale. We built a chip that's the size of a dinner plate, whereas most chips are the size of a postage stamp. We did that because we knew that this workload would go big and that the problems of memory bandwidth, problems of breaking up work and spreading it over lots of little machines would be daunting. This is my fifth startup and this is the first time I was wrong on the market size on the underside. I had no idea it was going to be this big. I think very few people saw, even those of us who were in it, how big this was going to be.
Early 2016, we went out, we did eight pitches, we got eight term sheets, we raised money, and we're building multi-ExaFLOP AI supercomputers for customers around the world.
**Sarah Guo** (3:27)
It's really incredible thinking about the foresight. I actually went back and looked at my notes from our end of 2015 meeting.
I'm sure you remember very well, but for our listeners, Andrew had a slide in there of how the top seven problems for deep learning were all long training times, and the whole industry was bottlenecked on computing. This was even before the scale up of Transformers that happened. It's kind of wild how correct you were, at least on the depth of this problem.
**Andrew Feldman** (3:59)
I think what you need to do is save that snippet and send it to my wife.
We got that right. I think also we got the fundamental architecture right in a large sense. We laid down the architecture before Transformers exist, and we're the fastest at Transformers by a lot. So when you break up a problem and it encounters things it has never seen before, and it's still really good at them, that's a sign in hardware architecture that you really got the architecture right.
**Elad Gil** (4:27)
What are some of the benchmarks that you use in order to assess the performance of your chip versus others, and how have you all performed relative to GPUs and other more standard?
**Andrew Feldman** (4:36)
We are not big believers in canned benchmarks. When I was at AMD, we had a team of 30 whose job it was to game benchmarks. When our CTO was at Sun, they had a team of 70 whose job it was to game benchmarks. How long it takes your customer to train a model is the answer.
22 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000627062735