**Jordan Nanos** (0:07)
Hello everyone, welcome back to episode number three of the Semi Analysis Weekly podcast. I'm Jordan Nanos, I'm here with Wega Chu, Myron Xie and Howie. The undisclosed Howie from an undisclosed location. Wega's also got his glasses on, he's ready to play poker.
And Myron's back, so anyway, this week we're going to talk about the article that we put out on Vera Rubin, extreme co-design, and we're going to walk through a lot of the updates that we're expecting this year. Let's talk a little bit about sparsity, peak marketed flops, the changes from grace to Vera, the system, server side, how everything's going cable-less, and yeah, I've got the three experts to talk about it here, all guys that are working on our accelerator and supply chain team for AI. So yeah, guys, welcome to the show.
**Howie** (1:10)
Good to be on.
**Wega Chu** (1:10)
Thank you, Jordan.
**Myron Xie** (1:11)
Thanks, Jordan.
**Jordan Nanos** (1:14)
All right, let's start by talking about like Vera Rubin and the concept of extreme co-design. Obviously, Rubin is the new GPU that's going to replace Blackwell. There have been a lot of improvements. Can you guys give me just a high-level overview of what's new with it, how you would introduce these changes?
**Myron Xie** (1:37)
Howie, if you bring up the Rubin floor plan here, we see a lot of interesting things. What's interesting with Rubin is we have, compared to Blackwell, it's two big compute die.
Eight stacks of HPM, but HPM moves to HPM4. And then also we have the IO chiplets being disaggregated into separate chiplets on the sides for the ME-Link C2C and the ME-Link 6 And then also with this comes a big improvement in memory bandwidth and also peak market and flops. For FP4, it goes to 35 petaflops dense versus 15 for Blackwell Ultra. And they also have a sparse 50 petaflops, which is interesting because this is a different type of sparsity. Howie, can you talk a bit about the adaptive compression engine that we see in Rubin and how it differs from the sparsity that Nvidia has marketed previously for the Blackwell and Hopper generations?
**Howie** (2:58)
Yeah, sure. So in this example, Nvidia is actually marketing up to five times increase in sparse FP4 performance compared to Blackwell GB200. But what they are actually comparing is sparse on Rubin versus dense on Blackwell. And why they are able to claim that is because of this new hardware compression in the transformer engine. So what it does is instead of doing structured sparsity, which is what has been done before, where you force every other data point in your tensor core, in your matrix into zeros, now the transformer engine directly looks at the data stream and dynamically compresses the data such that you can deliver up to 50 petaflops of effective compute while the actual processing is done at 35 So it's almost like a hybrid approach where you don't have the accuracy losses of forcing the data type to be zero, while also achieving some level of speed up. And why we believe this is important is because structure sparticity wasn't really used at all, right? Especially in the lower position down to FP4, the models wouldn't converge and there's a lot of accuracy issues. So now we believe with this new adaptive compression, you can effectively reclaim your sparse performance.
Yeah.
**Jordan Nanos** (4:38)
Yeah. So maybe take it back a step just to do some history. Like we've been dealing with Jensen math for four or five generations now since sparsity got introduced as a concept where people would double the peak theoretical flops for a given data type. FP32, then FP16 with the BFLOAT16 data type, and then we went to FP8. Now we're at the FP4 where there's different variations. Nvidia obviously pushing their NVFP4.
So the peak marketed flops of 50 petaflops is like a huge generation on generation increase that you might think is attributable to this Jensen math with sparsity. So to be clear, the difference with this one is that we expect this sparsity to work or to be useful in a way that previously it wasn't.
**Howie** (5:38)
Exactly.
**Jordan Nanos** (5:42)
Okay, cool. What else? Obviously, we got the picture on screen for those who are just listening.
The point about the 50 petaflops and the sparse NVFP4 points to the SM or the SM's within the GPC on the GPU, but there's plenty of other stuff that makes up this GPU. We got PCIe interfaces, we've got EV-Link C2C, EV-Link 6, we've got the HBM controllers with 288 gigs of HBM4. Maybe let's go through that one by one, starting with HBM. The two fundamental pieces here, FLOPs and then memory bandwidth. We don't actually have a spec, a firm spec on the HBM4, depending on the SKUs, but it's a significant increase over the Blackwell generation still.
34 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000751663520