Speed will win the AI computing battle  with Tuhin Srivastava from Baseten artwork

Speed will win the AI computing battle with Tuhin Srivastava from Baseten

No Priors: Artificial Intelligence | Technology | Startups

March 21, 2024

At a time when users are being asked to wait unthinkable seconds for AI products to generate art and answers, speed is what will win the battle heating up in AI computing.
Speakers: Sarah, Tuhin Srivastava, Elad
**Sarah** (0:05)
Hi, listeners. Welcome to another episode of No Priors. Today, Elad and I are catching up with Tuhin Srivastava, the CEO and co-founder of Baseten, which gives teams fast scalable AI infrastructure, starting with inference.
They're one of the players at the center of the battle heating up around AI computing. Welcome, Tuhin.

**Tuhin Srivastava** (0:22)
Hi. Thanks for having me.

**Sarah** (0:24)
Let's start at the beginning. For any listeners who don't know, what is Baseten and how do you start working on it?

**Tuhin Srivastava** (0:29)
Baseten is an infrastructure product, so we provide fast scalable AI infrastructure for engineering teams working with large models. Currently, we're focused on inference and we want to do a lot more after that. But for the past, say, four and a half years actually, that's a long time, for the last four and a half years, we've been cutting our teeth and trying to build this thing. I think it's been pretty rewarding over the last 12 months seeing the market show up and everyone get equally excited about AI infrastructure.
We started this honestly because firstly, we thought ML was pretty cool in 2019 We thought it was going somewhere.
We wanted to build a Pics and Shelfers business and solve the problems that we were running into. I think the side note here is that I want to start a company with my friends.

**Sarah** (1:15)
You often say that Baseten isn't no code, it's efficient code. Why does that difference matter?

**Tuhin Srivastava** (1:20)
That wasn't always the case, I'd say. I'd say there was times when we had elements which were definitely a bit no-code-y. I think what we've learned over the last three or four years is code is just incredibly powerful and engineers want to write code. Even in its best form, you want to build really, really tight abstractions, but I think the ability to turn the knobs under the hood is very, very important. I think no-code makes that a lot harder. I don't think it removes it, but it makes it a lot harder.
So what we do is just build very strong intuitive abstractions that try to make the easy things super easy and still make the hard things possible.
So you can get a lot of value really quickly, but I say, unlike a lot of other infrastructure products that have been built over the last 10 years, we're trying to solve against the graduation problem, which is that we're able to support teams as they grow in scale.

**Sarah** (2:15)
And just to sort of make it a little bit more visceral for our listeners, like what are the types of applications that run on Baseten? Like what's the scale of the platform? Do you have a favorite application?

**Tuhin Srivastava** (2:25)
Everything from tiny side projects on weekends, all the way to companies that are pretty AI native. We've supported foundation model companies. We work with companies like Descript, where AI is very, very cool to the product experience. We power a lot of AI features that Patreon has shipped. But I'd say some of the more interesting use cases from our perspective actually, or my perspective at least, are either the really small teams that we're giving a lot of leverage to, so that they can ship things very quickly. So a really good example of that might be a company like Planned AI, which is basically building an SDK for call centers, is how I describe it. But they're able to ship models and co-locate workloads so that they can get sub 300 millisecond or sub 200 millisecond responses without months and months of infrastructure effort. I think on the other hand, it's really exciting to see companies become AI enabled.
That's where we see a lot of the value is going to be over the next decade is, if I look at a company like Picnic Health, which has actually been around for a decade and starting to do very, very interesting thing with this corpus of data that they've gathered over the last 10 years and supporting those use cases, I think their models could Picnic GPT, which extracts information from medical records. And to me, those are the really exciting use cases where you're giving leverage to companies that are good at the domain that they are working in.
Their model might be proprietary, the data might be proprietary, but the infrastructure doesn't necessarily need to be proprietary and we can give them just an easy way to deploy that stuff without many, many people months.

**Sarah** (4:09)
It's become like in Vogue to compare the size of your GPU cluster, like people are spending a lot of money on GPUs. We hear about 600,000 H100 equivalents and lots of venture rounds being raised, often to train large models in some domain or another or even more and more expensive post-training.

30 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000649979369