**Skyler Monroe** (0:10)
Hey everyone, welcome back to the AI Hardware Show. I'm Skyler Monroe, your resident silicon obsessive, and we have a packed episode for you today. Huge thanks to our sponsors, AIEDA, helping businesses actually integrate AI into their real-world workflows. Ago Consulting, that's Ago, your go-to for silicon development from AI Accelerators to Full SOZ design, and Zen Semiconductor, an AI and Silicon Venture group building the processors and AI fabric behind modern data centers, with exciting ventures like their Sierra RISC-V CPUs and Loom AI Fabric. Today we're talking about AMD and Cerebras teaming up to take on Nvidia, we're looking at what a pair of Grace Blackwell chips can do sitting on your desk, and we're diving deep into Nvidia's next-gen Vera Rubin architecture. Let's get into it. Alright, so story one, and this one is genuinely exciting if you follow the inferencing wars AMD and Cerebras have just announced a partnership that directly counters Nvidia's acquisition of Groq. And I want to walk you through why this matters, because there's a lot of layers here. So first, let's set the stage.
Nvidia picked up Groq, well, the LPU technology recognizing that inference is the next massive battleground in AI compute. Training gets all the headlines, but inference is where the real money is long-term. Every time you type a prompt into an AI system, that's inference happening at scale, and it needs to be fast and efficient.
Now, Groq's language processing unit, the LPU, is a really clever piece of engineering. Instead of using a traditional GPU approach, it clusters hundreds or even thousands of specialized chips together. Each individual chip packs giant blocks of what they call matrix multiply units, the MXM blocks and vector units, the VXM blocks plus around 230 megabytes of blazing fast on chip SRAM. The key insight there is that SRAM, it's dramatically faster than HBM or GDDR memory, but it's also way more expensive and power hungry per byte. So Groq's bet is essentially keep the model weights resident in super fast memory across a massive cluster, and you get incredible token throughput with minimal latency. It's like having a library where every book is already open on the table, rather than having to pull them off shelves one by one.
Now, Cerebras takes a completely different and honestly kind of wild approach. They build the wafer scale engine, which is exactly what it sounds like. Instead of cutting a silicon wafer into hundreds of individual chips, Cerebras just doesn't. They use the entire wafer as one massive chip. We're talking about something 58 times larger than a conventional GPU die. That size advantage translates into an enormous amount of on-chip compute, and crucially, on-chip memory bandwidth. When you're that big, you can connect compute units with very short, very wide on-chip interconnects rather than going off-chip through slower memory buses. The result, according to the numbers being thrown around, is up to 15 times faster than Nvidia's GPS for certain workloads. So, where does AMD fit into all of this? AMD has developed something called Helios, which is their rack-scale networking and integration solution. Think of Helios as the connective tissue that ties together, compute at the rack level, managing data movement, memory coherence and interconnect between accelerators at scale.
By fusing Helios with the Cerebras wafer scale engine, AMD is essentially building a rack-scale inferencing beast.
The partnership reportedly achieves 5 times higher tokens per second per watt compared to alternatives. That tokens per second per watt metric is really the holy grail for inferencing right now. It's the efficiency figure that determines your cost per query at scale. Think about it from a data center operator's perspective. If you're running a billion queries a day, that efficiency multiplier directly translates to your electricity bill and your cooling infrastructure costs.
5 times better efficiency isn't incremental, that's a generational leap. What I find fascinating about this pairing is that it's a classic, the enemy of my enemy is my friend situation. And actually, the register used exactly that headline, which I love. AMD has GPU assets and rack-scale integration expertise through Helios. Cerebras has the most radical, highest-performance inference silicon on the market. Neither one can beat Nvidia alone right now, but together they're attacking Nvidia's newly acquired inference crowned from a very credible angle. The question going forward is software ecosystem. Groq has been building out its developer tools and Nvidia obviously has the CUDA moat. AMD and Cerebras will need to make this integration dead simple for developers to actually win customers. But the hardware story here? Genuinely compelling. Now, zooming out a bit, this also signals that the AI inference market is fragmenting. You're not going to see one architecture win everything. Different use cases, latency-sensitive chat, batch processing, RAG pipelines will favor different hardware, and that's actually healthy for innovation. The old Just Buy H100's playbook is getting more complicated by the month. Alright, let's pivot to story 2 And this one is more of a lifestyle piece, which is a phrase I never thought I'd say about enterprise compute hardware. But here we are. Tom's Hardware tested what happens when you pair two Dell Pro Max systems, each running Nvidia's Grace Blackwell GB10 chip and connect them together into a small local cluster. And the headline number that jumps out is the price. Each unit runs about $6,332.
8 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000778151288