**Anjney Midha** (0:00)
It was one of our founders who came up with the name Oxygen because they basically said, look, if I don't have that kind of compute on day one, I can't breathe. So on day one, we were able to then say to founders, look, you have guaranteed capacity at prices you just can't get anywhere else. While saying to our cloud compute partner, look, you get direct access to the world's best foundation model startups and AI startups. They realize the value in that. For the founders, it's very clear what they get. They're able to raise less and take on less long-term risk while still being able to train really great models on day one. But our goal is always to try to give start-ups unfair advantages compared to big tech companies. Just by resetting compute to rational sort of normal market rates, we were able to give these teams an unfair advantage.
**Derrick Harris** (0:42)
Hi, you're listening to the a16z AI podcast. I'm Derek Harris, and joining me once again this week is a16z general partner Anjney Midha, this time to discuss the economics of GPUs for AI workloads, and a program a16z is running to help companies in our portfolio access them at a reasonable price. I promise it's not an advertisement for that program, which is called Oxygen, but more discussion about how it came to be so difficult for startups and other small customers to acquire adequate infrastructure from cloud providers. One very interesting insight for folks who aren't entrenched in the world of AI and cloud economics is how we have, as Anjney puts it, gone back to the future in terms of the capital expenditure required to launch a startup, whereas infrastructure as a service was supposed to let startups avoid the overhead of buying new servers and the high risk of over-provisioning. Cloud provider requirements for long-term contracts, paired with bidding wars from deep-pocketed AI incumbents, have essentially put startups trying to train foundation models back in that same pre-cloud situation. If they must commit to three years and many millions of dollars, they need incredible customer demand in short order or they're left sitting on and paying for very expensive compute capacity that they don't need. If you've heard about other investors supplying affordable GPUs as a value add to startups, this is why. But you've probably never heard the rationale behind these efforts explained in such detail.
As a reminder, please note that the content here is for informational purposes only, should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any a16z fund. For more details, please see a16z.com/disclosures.
**Anjney Midha** (2:30)
So Oxygen is our compute program at a16z, where we help startup founders and our companies navigate their compute challenges. Whether it's helping them find the capacity they need for training, whether it's for inference, we have a number of options now for startup founders, particularly those working on large-scale AI infrastructure efforts, who might have very capital-intensive, GPU-hungry business plans to be able to access the kind of compute they need in a timely way with our help. As the scaling laws in AI were becoming more and more mature, it just started to be hard to ignore just how much of my time I was spending helping founders just navigate their compute needs. It started with Anthropic, I would say, in early 2021, when I got a call from Dario and Tom, two of the co-founders of Anthropic, and had been leading the GPT-3 efforts at OpenAI. And around the time they had decided to leave and start Anthropic, they gave me a call and said, hey, we'd love to get you involved as an early investor. And I said, sure, you know, what are you thinking of raising for your seed round? And they said, we need 500 million to get started.
And that was a bit of a shock. And soon after that, I started realizing that their needs weren't isolated. So I think this started from a working backwards kind of realization that a number of the customers we serve every day, which are founders, especially working at the frontier of AI infrastructure, all had a common problem, which was as a startup, they were being de-prioritized by the large GPU clouds, the hyperscalers, in favor of larger customers, which is really tough. We were in the middle of a supply crunch at the time where H100 capacity was in short supply. And as a result, what was happening was the hyperscalers who run cloud businesses that have very sensitive margins tied to their occupancy rate or their utilization rate for their clusters were basically starting to prioritize long-term contracts over short-term contracts, which is totally the rational thing to do. But if you're a startup and now for you to access the same price for hourly GPUs that you could have gotten just six months ago for a six-month, call it rental contract, if you had to now buy a three-year contract, you're often being asked to commit by the hyperscalers more capital than you'd raised or even plan to raise in the next year upfront to get access to those rates. And so illustratively, what was happening at the time was the market rate for short-term GPU capacity, 3 to 4x over that period, I would say, of late 2020 to mid 2023 That was a real, I would say, a realization moment where it wasn't like there were some mass conspiracy theory or anything against startups, but just the natural market forces made it such that if you were a startup in foundation model land, and you wanted to be able to get access to any significant number of GPUs on day one, it was extremely hard for you to do that at, call it sane and rational prices without committing to a two, three, four-year sometimes commitment on those GPUs. Now, that's a really hard thing for you to do as a startup for three reasons. One is, early on, you haven't raised that much capital, and so it's very daunting to commit more capital than you've even raised. The second thing it does is, it makes it very difficult to do capacity planning when you don't even know what your inference needs are going to be like. It forces you to have to make a bunch of suboptimal decisions about your capacity. In the normal scheme of things, how this would work is you get started as a startup, you go buy some short-term capacity for, let's say, six months. You then need to train your foundation model over six months. You then release the model, you start getting customers, and you have a pretty good sense at that point of your demand from your customers for inference. You know which days of the week it spikes. You understand which regions you're getting the most inference demand from. You understand what the queue times are like when you release new features, and then you use that to inform your purchasing for inference.
33 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000674160258