Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud artwork

Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud

No Priors: Artificial Intelligence | Technology | Startups

May 1, 2026

Baseten CEO and co-founder Tuhin Srivastava sits down with Sarah Guo and Elad Gil to discuss the rapid growth of AI inference demand, Baseten’s 30x growth, and why inference is becoming the strategic “last market.
Speakers: Sarah Guo, Elad Gil, Tuhin Srivastava
**Sarah Guo** (0:06)
Today Elad and I are here with Tuhin Srivastava, the founder and CEO of Baseten, the AI Inference Cloud. We're here to talk about capacity constraints for AI compute, why inference is the last market, how the workload is changing, the open source and perhaps multi-chip future, and what 30x scale in a year looks like. Tuhin, welcome back.

**Elad Gil** (0:28)
Hi. Good to see you. Thanks for having me.
All right.

**Sarah Guo** (0:31)
You are in one of the craziest markets, AI Inference. It's very important. There's a lot going on. You guys have grown 30x over the last year.
I think I can say you're expecting to do more than a billion dollars in revenue this year. What's going on? Tell us about scale.

**Tuhin Srivastava** (0:49)
Yeah. No, it's been nuts. I think what's happened over the last, I don't see it 24 months, but it just keeps getting bigger and bigger is that, I think everyone is realizing that you can put AI everywhere.
You have all these great options available from closed source to open source models. The open source models have crossed some chasm in terms of their baseline capability. I think RL techniques and post-training for specialized models has become main treatment of them. There's enough examples of it working, the customers realizing they can own their inference more and more. What that's meant for us is more the long tail of models coming true, customers in-housing a lot of that intelligence themselves.
As the application layer just gets bigger and bigger and bigger and that's growing, we are just someone index on that and we've been around to be able to collect the demand.

**Sarah Guo** (1:55)
There's an existential question in here that I think everybody is continually asking of, does the independent application layer get to exist at all versus the labs? You have to believe this, why do you believe it?

**Tuhin Srivastava** (2:07)
Yeah, look, I think it would be a sad thing if it didn't exist in general, but you know, sadness is fine.

**Elad Gil** (2:16)
Sad all the time.

**Tuhin Srivastava** (2:17)
Sadness is fine, but that's not the reason why I think the application layer will exist. I think the application layer will exist for a number of reasons. One is because I think this idea that what is valuable to a company is the user signal that they can gather, that only they can gather.
And to the extent that that is encoded in a model, I think a lot of their business will be at risk, but to the extent that it is encoded in workflows, that is where they will be able to develop notes. So a good example of that is, say, a company like Abridge, where the clinicians edits all the notes and what they do with those notes after the fact, and the thing that happens inside the EMR three steps down, that becomes a workflow that only...

**Sarah Guo** (3:11)
Can you explain what Abridge does?

**Tuhin Srivastava** (3:13)
Sorry, Abridge is an ambient scribe that is used by physicians in...
You know, almost all hospitals in the US. I think Elad's an investor, Shiv's amazing, great company, great team, great product.
And they've basically got this very, very deep integration into hospitals, into clinician workflows. And my argument would be here is that actually, it's very, very hard for a frontier model company to be able to eat up away at that because they just don't have access to that user signal. And what will happen over time is the folks who have access to that user signal can start to post-train models on that reward signal and start to get long-horizontal models running that. And I think to the extent that that is possible and that signal is differentiated and unique and is somewhat rare to get access to, there will be an application layer. And I think support companies is another example of that, where a support task isn't one-shotted. Usually at a company like Baseten, when a ticket comes in, there's like one, two, ten, twenty actions that get taken. And that is where someone can develop a specialized model.

**Elad Gil** (4:34)
So there's almost two versions of this then. There's the new companies like Abridge or Dekagon or some of these other things that you mentioned that are doing these new types of applications that are using AI and they sell it to customers. The other is enterprises building things in-house or building their own models. What proportion of the market today do you think is these new application companies versus enterprises adopting AI?
And how do you think that looks in a couple of years?

35 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000765654238