**Swyx** (0:03)
Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before. As well as Ali, welcome.
**Ali Taha** (0:15)
Pleasure to meet you.
**Swyx** (0:15)
Waterloo intern.
**Ali Taha** (0:16)
Waterloo intern, always.
**Swyx** (0:17)
When did you get Waterloo intern as a handle?
**Ali Taha** (0:20)
I think the rebranding happened like mid-March. When I saw it was open, I was like, I have to take it for grabs.
**Philip Kiely** (0:25)
The problem is that Ali is really good at his job, and he's not going to be an intern much longer. So we have to figure out who's going to get the handle.
**Ali Taha** (0:33)
Pass the torch over to Ali.
**Swyx** (0:35)
You can just pass it to another Waterloo intern.
**Ali Taha** (0:37)
Another Waterloo intern. You gotta get an intern from Waterloo. You gotta get an intern from Waterloo.
**Swyx** (0:45)
But it could come from Baseten, so it's like whoever Baseten gets from Waterloo has the title of Waterloo.
**Vibhu** (0:50)
They have to pass it to Ali. Halfway through, you either get it or you're out.
**Swyx** (0:55)
You should also do a big graduation ceremony where you change the handle.
**Vibhu** (1:00)
I mean, you guys are good at ceremonies clearly. We had a nice launch of the book, very successful. But before we get into all that, I want to start off with a fun question for you.
You're an expert inference engineer. What happens when I send a long query, say 200,000 tokens into Baseten's inference? What's the process of query through GPU, model, routing, balancing, all that? What is all the stuff that we don't think about?
**Philip Kiely** (1:25)
With a long query specifically, the first thing that I'm going to ask is, have you sent me this query before or at least part of it? I really hope you have because it's going to be a lot easier for me and a lot cheaper for you.
The first thing that we're going to look at is some cache-aware routing where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with number one, available pre-fill workers and number two, ideally some cached input already there so that we can skip pre-fill on at least part of these 200,000 tokens. If you're doing 200,000 tokens, it's probably coding or a multi-tone agent or something where you would expect to have that cached.
If you don't, we're going to have to send it to a pre-fill worker. We've, at least on certain models, disaggregated pre-fill and decode. So you're going to have one set of GPUs that's solely going to process the input, create that KV cache, and get you your first token. And then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some kind of speculator model in front of that. I'm going to assume that you're doing coding. And because of that, our speculator model, which assumes you're doing coding, is going to have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's going to be slower. And then we stream that output to you and account for it, charge you some number of couple of pennies and say, hey, would you like to send another one?
**Swyx** (3:04)
Except Baseten isn't charged by pennies.
**Philip Kiely** (3:07)
Well, yeah, we charge... I'm assuming that we're talking about the public model APIs.
If you are setting up a dedicated deployment, then yeah, it's not pennies.
**Swyx** (3:18)
Yeah, I mean, one of the key differentiators when I was talking with Basen initially was that actually people who want very, very high volume just need to rent by the box because then it's up to you to figure out how to saturate the box.
**Ali Taha** (3:31)
And more often than not, it's like way cheaper if you're pushing like millions of tokens per hour, if you just pay per hour instead of paper token.
**Philip Kiely** (3:37)
Yeah, they do. I think that we've increasingly seen a lot of demand for the sort of paper token API is just because everyone wants to try open models. And then once they find a use case that's really sticky, then they move over to dedicate it.
**Vibhu** (3:51)
Is there a best practice on when it's time to swap over?
**Philip Kiely** (3:54)
A couple of reasons. Yeah, reliability, that's a big one, right?
96 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID