**Alessio** (0:05)
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel, and I'm joined by Michael Swyx, founder of Smol AI.
**Swyx** (0:13)
Hey, and today we're so excited to be finally in the studio with Evan Conrad from SF Compute.
**Evan Conrad** (0:17)
Welcome. Hello, how goes it? How are we doing?
**Swyx** (0:20)
I've been fortunate enough to be your friend before you're famous, and also we've hung out at various social things. So it's really cool to see that SF Compute is coming into its own thing, and it's a significant presence, at least in the San Francisco community, which, of course, is in the name, so you couldn't help but be.
**Evan Conrad** (0:38)
Indeed. I think we have a long way to go, but thanks.
**Swyx** (0:41)
Of course. One way I was thinking about kicking out this conversation is we will likely release this right after Coreweave IPO.
And I was looking, doing some research on you. You did a talk at The Curve. I think I may have been viewer number 70 It was a great talk. More people should go see it, Evan Conrad at The Curve, but we have like three orders of magnitude more people, and I just wanted to highlight like, what is your analysis of what Coreweave did that went so right for them?
**Evan Conrad** (1:11)
Sell locked in long-term contracts and don't really do much short-term at all. I think like a lot of people had this assumption that GPUs would work a lot like CPUs. And the like standard business model of any sort of CPU cloud is you buy commodity hardware, then you lay on services that are mostly software, and that gives you high margins.
And pretty much all your value comes from those services, not really the underlying compute in any capacity. And because it's commodity hardware and it's not actually that expensive, most of that can be sort of on-demand compute. And while you do want locked in contracts for folks, it's mostly just a sort of de-risk situation. It helps you plan revenue because you don't know if people are going to scale up or down. But fundamentally, people are like buying hourly, and that's how your business is structured. And you're going to make 50% margins or higher. This doesn't really work in GPUs. And the reason why it doesn't work is because you end up with super price-sensitive customers. And that isn't because necessarily it's just way more expensive, though that's totally the case. So in a CPU cloud, you might have, let's say if you had a million dollars of hardware, in GPUs you have a billion dollars of hardware. And so your customers are buying at much higher volumes than you'd otherwise expect. And it's also smaller customers who are buying at higher amounts of volumes, so relative to what they're spending in general. But in GPUs in particular, your customer cares about the scaling law behind it. So if you take like Gusto, for example, or Rippling or an HR service like this, when they're buying from an AWS or a GCP, they're buying CPUs and they're running web servers. Those web servers, they kind of buy up to the capacity that they need. They buy enough CPUs and then they don't buy any more. Like, they don't buy any more at all.
**Swyx** (2:55)
Yeah, you have a chart that goes like this and then flats.
**Evan Conrad** (2:57)
Correct, and it's like a complete flat. It's not even like an incremental tiny amount. It's not like you could just like turn on some more nodes and then suddenly, you know, they would make incremental amounts of money more. Like, Gusto isn't going to make like, you know, 5% more money. They're going to make zero, like literally zero money from every incremental GPU or CPU after a certain point. This is not the case for anyone who is training Modals, and it's not the case for anyone who's doing test time inference, or like inference that has scales at test time. Because like, your scaling laws mean that you may have some diminishing returns, but there's always returns. Adding GPUs always means your model does actually get better, and that actually does translate into revenue for you. And then for test time inference, you actually can just like run the inference longer and get a better performance, or maybe you can run more customers faster and then charge for that. But it actually just translate into revenue. Every incremental GPU translates to revenue. And what that means from the customer's perspective is you've got like a flat budget, and you're trying to max the amount of GPUs you have for that budget. That's very distinctly different than like where Augusto or Rippling might think, where they think, oh, we need this amount of CPUs. How do we reduce our amount of money that we're spending on this to get the same amount of CPUs? What that translates to is customers who are spending in really high volume, but also customers who are super price-sensitive, who don't give a shit, can I swear on this? Can I share our sports? Who don't give a shit at all about your software, because a 10% difference in a billion dollars of hardware is like a hundred million dollars of value for you. So if you have a 10% margin increase, because you have great software, on your billion dollars, the customers are that price-sensitive. They will immediately switch off if they can, because why wouldn't you? You would just take that hundred million dollars, you'd spend 50 million dollars on hiring a software engineering team to replicate anything that you possibly did. So that means that the best way to make money in GPUs was to do basically exactly what Coreweave did, which is go out and sign only long-term contracts. Pretty much ignore the bottom end of the market completely, and then maximize your long-term contracts with customers who don't have credit risk, who won't sue you if or are unlikely to sue you for frivolous reasons. Then because they don't have credit risk and they won't sue you for frivolous reasons, you can go back to your lender and you can say, look, this is a really low risk situation for us to do. You should give me prime interest rate. You should give me the lowest cost of capital you possibly can. When you do that, you just make tons of money. The problem that I think lots of people are going to talk about with Coreweave, is it doesn't really look like a cloud provider financially.
72 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000748427955