**Sarah** (0:05)
Hi, listeners. welcome to No Priors. Today we're talking to Jared Quincy Davis, the founder and CEO of Foundry. Jared worked at DeepMind and was doing his PhD with Matei Zahari at Stanford before he began his mission to orchestrate compute with Foundry. We're excited to have him on to talk about GPUs and the future of the Cloud. welcome, Jared.
**Jared Quincy Davis** (0:23)
Thanks, Sarah, and great to see you. Thanks a lot as well.
**Elad** (0:26)
Yeah, great seeing you.
**Sarah** (0:27)
The mission at Foundry is directly related to some problems that you had seen in research and at DeepMind. Can you talk a little bit about the genesis?
**Jared Quincy Davis** (0:35)
A couple of the most inspiring events I've witnessed in my career so far were the release of AlphaFold 2 and also Sedge of Tea. I think that one of the things that was so remarkable to me about AlphaFold 2 is initially it was a really small team, three and then later 18 people or so, and they solved what was kind of a 50-year grand challenge in biology, which is a pretty remarkable fact that every university, every pharma company hadn't solved. And similarly with TechPT, a pretty small team, opening out with 400 people at the time, released a system that really shook up the entire global business landscape. That's a pretty remarkable thing, and I think it's kind of intriguing to think about what would need to happen for those types of events to be a lot more common in the world. And although those events are really amazing because of the small numbers of people working on them, I think it's not quite the David and Goliath story, neither are quite the David and Goliath story that they appear to be when you when you double click. In OpenAI's case, they were only 400 people but had 13 billion dollars worth of compute, which is quite a bit of computational scale there. And in DeepMind's case, it was a small team, but obviously they were standing on the shoulders of giants in some sense with Google, right? And the leverage that they had via Google. And so one thing I think that we thought about is, what can we do to make the type of computational leverage and tools that are currently exclusively the domain of OpenAI and DeepMind kind of available to a much broader class of people? And so that's a lot of what we worked on with Foundry, saying, can we build a public cloud? You build specifically for AI workloads, where we re-imagine a lot of the components that constitute the cloud end-to-end from first principles. And in doing that, can we make things that currently cost a billion dollars cost 100 million and then 10 million over time? And that'd be a pretty massive contribution. I think it would increase the frequency of events like AlphaFold 2 by 10x, 100x, or maybe even more, super linearly. And we're already starting to see the early signs of that. But quite a lot of room left to push this agenda. So really exciting. So that's kind of maybe an initial introduction, preamble, to how we thought about it. And I can trace that line of reasoning a bit more, but that's kind of part of what we've done.
**Sarah** (2:41)
Jared, for anybody who hasn't heard of Foundry yet, what is the product offering? Yeah.
**Jared Quincy Davis** (2:47)
So Foundry, we're essentially a public cloud built specifically for AI. And what we've tried to do is really reimagine all of the systems undergirding what we call the cloud, end-to-end, from first principles for our workloads. And we've tried to do this a bit of a new way. I think the AI offerings from the existing major public clouds and kind of some new GPU clouds, haven't really re-envisioned things. And by thinking about a lot of these things a bit anew, we've been able to improve the economics by, in many cases, 12 to 20x over lower tech GPU clouds and existing public clouds. And we'll partially base on some of these products that we'll talk about today that we're releasing and a lot of new things that we're working on. We think we can push that code a bit further as well. And so, our primary products are essentially infrastructure as a service, so our customers come to us for elastic and really economically viable access to state-of-the-art systems. And also a lot of tools to make leveraging those systems really seamless and easy. And we've invested quite a bit in things like reliability, security, elasticity, and just the core price performance.
**Elad** (3:56)
How underutilized are most GPU clouds today? And I think there's almost three versions of that. There's things on hyperscalers like AWS or Azure. There's large clusters or clouds that people who are doing large scale model training or inference run for themselves. And then there's more just like everything else. It could be a hobbyist. It could be a research lab. It could be somebody with just some GPUs that they're missing around with. I'm sort of curious for each one of those types of or categories of users, like what is the likely utilization rate and how much more do you think could be optimized? Is it 10%? Is it 50%? Like I'm just very curious.
39 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000666226720