Ep. 010 - How Much Do GPUs Really Cost, and Where Does the Value Go? (AI Cloud TCO) | Jordan Nanos, Dan Nishball, Kang Wen Cheang, Zane Fong artwork

Ep. 010 - How Much Do GPUs Really Cost, and Where Does the Value Go? (AI Cloud TCO) | Jordan Nanos, Dan Nishball, Kang Wen Cheang, Zane Fong

SemiAnalysis Weekly

May 1, 2026

This episode features Jordan Nanos (@JordanNanos) and Daniel Nishball (@dnishball) breaking down the economics of GPU clusters through real-world data and experience.
Speakers: Jordan Nanos, Daniel Nishball, Kang Wen Cheang, Zane Fong
**Jordan Nanos** (0:03)
Well, I saw Akash's, Mike, I was reminded of this Vine. Dan, do you remember Vine before TikTok?

**Daniel Nishball** (0:10)
Yeah, that died, yeah, I remember that.

**Jordan Nanos** (0:13)
Do you remember those ones of the guys who would have the room fan going, and they'd sing into the fan, and it'd sound like auto-tune? Baby girl, what's your name?

**Daniel Nishball** (0:25)
Shorty, Shorty come downtown.

**Jordan Nanos** (0:32)
Kang Wen, this is before your time, dude.

**Daniel Nishball** (0:35)
Yeah, Kang Wen doesn't even know what Shorty is.

**Kang Wen Cheang** (0:37)
Who the f*** is Shorty?

**Daniel Nishball** (0:40)
Hi, everyone. Welcome to SemiAnalysis Weekly. Jordan and I are switching roles at the turns of table this week.
I'm going to be hosting Jordan as he talks about a new article that we've done on rethinking the total cost of the GPU cluster. Kang Wen and I spend a lot of our time working on the total cost of ownership, but oftentimes, this is from a theoretical or quoted perspective. It's a bit like being at a car dealership and you're saying, this thing does 36 miles per gallon, but you know what? When you take it on the road, you have your mix of highway versus city, the performance is going to be different. That's what this article is all about. Jordan's had a lot of experience actually using GPUs on many clusters. He's worked with dozens, maybe even a hundred Neo clouds, and so he's drawn on that experience in terms of saying, hey, what is the actual on-the-road performance? Jordan, do you want to kick us off and talk about a few points in your article?

**Jordan Nanos** (1:32)
Sure. Yeah, thanks, Ben. I think the high level is in the title, which is how much do GPU clusters really cost. The article really explores implied costs that are not explicitly called out on a bill of materials or on a purchase order, but maybe just take a step back. When people are renting GPUs, in many cases, the price is not the same, and so you get a cheaper price from many different providers, but you really want to look beyond the individual price of an individual GPU and figure out how much useful work you can get out of the cluster. What we did in this article is introduce a framework for actually calculating good put. Good put being measured as how much useful work you can do on the cluster, how many jobs you can run, or how much inference you can serve. We got on a screen for the people viewing. Basically, we tried to include line items that people would explicitly purchase, which is the GPUs, any orchestration premiums, the cost of storage, the cost of networking, the cost of support nodes, these are CPUs. The cost of support itself, many of the hyperscalers charge a premium just to give you response times within a certain SLA. You got to pay more money if you want quicker responses basically. Then there's these implied costs, which is the time spent setting things up or running POCs, the time spent debugging on the cluster if there's issues, and then goodput, which is that basically way to calculate how much downtime you're going to have on the cluster while you operate it. As Dan said, it's like mileage on a car. If some of your GPUs are not accessible and they're not working and there's interruptions or failures while you're trying to run jobs, this doesn't just cost you time on that one GPU, it actually costs you time on the entire cluster because nothing can do useful work if one GPU fails depending on how you're setting up your clusters. Anyway, the article explores all of that.
There's lots of ways in which we can go from here to talk about maybe what sort of contribution we did on the research and what we found, which is maybe the high-level findings is that there's a big discrepancy, especially for big jobs on big clusters between top-tier providers that have less failures and a quicker recovery time from a failure when compared to lower-tier providers that just have a bunch of failures and it takes them a long time to recover from, which is previously to this, I think a lot of people didn't have a way to put numbers to that intuition. They'd know it would be somewhere 5, 10, 20 percent, I'd be willing to pay a premium, but exactly how much should you be willing to pay? We think we have a way of answering that based on this calculator, which for what it's worth is available for free for people listening at clustermax.ai.tco.

41 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000765663974