**Sam Whitmore** (0:00)
And you have an agent running right now.
**Charlie O'Neill** (0:02)
I do.
**Sam Whitmore** (0:02)
If I know from setup. What is your agent doing today?
**Charlie O'Neill** (0:06)
Today, as Harry mentioned, we've been doing some KBcache compaction stuff recently.
So trying to figure out how we can extend context windows by compacting the KBcache. So we have probably between us about 64 to 128 agents working on this at a given time.
**Sam Whitmore** (0:24)
Right now, how many guys are on that computer?
**Charlie O'Neill** (0:27)
Actively running jobs right now. I think I've got 16 nodes of HFUs each, and then the agents partition is how they like.
**Harry Partridge** (0:36)
I find a lot of the work, if you want to have a lot of agents going, I feel like normally I have a few agents that I actually talk to, and the other ones are delegated tasks. Like I tell my main agent to delegate tasks to the other ones. I set up some messaging scripts so that they can send direct user messages to each other.
And so, yeah, I'll just say like how's my... I gave them all like different like mathematician names, you know, like I'll be like, oh, how's Poincare doing today? Or, you know, what's Hilbert up to? You know?
**Sam Whitmore** (1:10)
That's cool. So you actually remember which one is doing what by the mathematician?
**Harry Partridge** (1:14)
Yeah, yeah, I just have like, oh, you know, Hilbert's working on the evals, you know?
**Sam Whitmore** (1:28)
I was very excited to meet you guys, and eager to hear maybe kind of what you work on at Baseten, and kind of how you got where you were. And to introduce myself, I'm Sam Whitmore. I work at Cursor. I am an engineer on the Cloud Agents team. I've worked at Cursor for about six months.
But yeah, hangover with you guys.
**Charlie O'Neill** (1:46)
So my co-founder is Moody Maxima. I started Puzz, I started last year.
And at the beginning, it was sort of blue sky research. I'm like, okay, we're very interested in open source models. We're interested in supporting an ecosystem that wasn't just one or two closed frontier models. How do we go about that? And I think we pretty quickly narrowed in on the thesis of specialization and post-training in particular as a way to allow people to own their own intelligence. And I think when we first started, that was quite a contrarian take. People were quite skeptical that you could really do economically valuable things with open source models. But we believed in the thesis, and I think sort of by the middle of last year, the base intelligence and capabilities of open source had gotten good enough that we could actually start to specialize them, particularly for these kind of sub-tasks or very repeatable things that a lot of companies were doing. And then, yeah, we were using Baseten for inference. We thought they were the best inference providers in the world. We tried all of them.
And then I think at one point, Baseten noticed that we were sort of directing a lot of new inference demand onto Baseten GPUs. And so we had a bunch of chats and realized that this would be a much more powerful thing if we could do this together, if you could couple the inference and the real world signals you were getting from inference and users telling you what they did or didn't like about the particular way that your language model was performing in your harness. And when you couple that with training, that's a really powerful paradigm. It kind of connects together this feedback loop, which is much more powerful than what you can achieve with prompting and constantly just iterating a new harness. And so yeah, we've been at Baseton ever since, and Harry was one of our first employees at Parviz.
**Harry Partridge** (3:30)
I wanted to work in AI, and I'd been interested in it for such a long time. And so, yeah, when I heard Charlie was doing a startup, you know, I had a chat, and yeah, just, you know, I was like, oh wow, Charlie really knows what he's talking about.
And so, so yeah, then I joined up and we had a lot of fun. I think the thing is because of the scale we were at, initially, you know, we really had to sort of scrap around and like figure out how to do things more efficiently. And so that was kind of fun. Yeah, like just, you know, kind of inventing some new sort of strategies that's going to like basically get us like really high good performance, but just, you know, without using like a ton of compute. Since then, I've been like, now we do have access to a ton of compute and it's a lot more fun. We can do a lot more larger scale experiments. So yes, it's great.
40 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000772251852