**Grace Gong** (0:02)
Hi Kai, welcome to Venture with Grace.
**Kai Mak** (0:04)
Hey Grace, thanks for having me. How are you?
**Grace Gong** (0:07)
I'm good. It's been a week. I was at the, as I was telling you, like right before this, I was at the AIE Conference, and obviously I saw the Together AI booth, and congrats on the Series C.
**Kai Mak** (0:21)
Thank you. Yeah.
**Grace Gong** (0:23)
So maybe we'll start from that.
So obviously, I think Together AI has been evolved many times, but in a really short amount of time, that you guys stress raised 350 last year, and then this is like an 800-milliliter run. Maybe we could talk about, obviously, I'm curious how you guys are going to spend the money, but let's start with a little bit background on Together AI, and how is thought raising going to shift towards the future?
**Kai Mak** (0:53)
Yeah, so I posted this on LinkedIn, actually, but I joined Together AI around two years ago.
I got to know Vipal. Vipal is just an incredible entrepreneur, both technical and incredibly business savvy. But when I first met him, there were two things you really had to believe. I started back in summer of 2024 You had to believe that open-source AI was going to be really, really big and it was going to be at parity with the frontier eventually. You also had to believe that inference was going to be this enormous workload. So if you looked like two years ago, you had to squint and you're like, okay, I think this is going to happen. Lama 3.1 had kind of just come out.
You fast forward to today and together AI, we built this AI native cloud really servicing and native companies for training, post-training, inference, basically the whole lifecycle. You fast forward to today and now tokens have completely taken off. Inference is the biggest workload, and it is going to grow orders of magnitude from here. Back to the fundraise, it is really to help continue to support the growth of our team, our infrastructure, and of course, our customers and partners as we go and help fulfill these tokens globally.
**Grace Gong** (2:38)
For sure. Maybe we could start with the inference piece because obviously I feel like it's very popular in the tech industry, but maybe we could talk about how does it shift how people were training model, and maybe we could start talking about this particular sector that you feel like will have the biggest impact with the fundraise. I guess where are you guys focusing on?
**Kai Mak** (3:06)
Yeah, maybe if I had spoken to other folks outside of AI previously, they would have thought that inference is easy, will be a commodity. Certainly, there's a bunch of open-source libraries out there to help companies with inference. But in fact, inference is very, very difficult. It is very hard, it is very bespoke, especially for companies using a lot of tokens based on the model they're using, as well as traffic shape.
We spend a lot of time and a lot of focus optimizing, providing a lot of thoughtful managed services for these companies using a lot of tokens.
As we think about the companies that we're working with, and then we will continue to focus on, especially after this fundraise, there are two distinct ideal customer profiles that we think deeply about. The first is certainly AI native companies. This is the Cursors of the World, the Decacons of the World. These are the folks who have just seen tremendous adoption, are thinking deeply about their unit economics as they are serving their customers. In doing so, working with a company like Together AI, where they can fine-tune their models, they can serve it with the best price performance such that their unit economics and gross margins make sense as they scale, I think will continue to be a growing sector, and so quite important to us.
The big reason why we love AI natives is because they are at the forefront of the types of services that they're providing their customers, and they're continually pushing us to think about the products and services that we want to provide on our Together AI Cloud.
That's one, and that will continue to grow enormously. Second is, since December of last year, now enterprises and digital natives are convinced to use AI for coding. There's a bunch of coding agents out there to do long-horizon tasks and do these tasks end-to-end. This is the most critical breakthrough. Obviously, you've seen the ARR, these frontier model labs, just go parabolic. Yeah, obviously, there's this whole, the talk back then was all around token maxing, and how do you make your engineers as productive as possible, and you fast-forward a few months, and quickly these digital natives are quite rational, and they're like, this doesn't make any sense. These tokens are out of control. There are some companies that they spent their annual token budget in the first three or four months of the year, right? And so, the second ideal customer that we're working with, they are, they have a bunch of software engineers, they're digital native companies. They, of course, are going to use AI and AI agents for coding, but they need to think about it in a rational way, such that it is productive, they have cost controls. And so, this is a company segment that we think we can provide tremendous value to. We've worked with several digital natives already who come to us to reduce their token cost by 95% from the frontier models, with basically app parity, because these open-source models are so good. And so, basically, we are working with AI native companies to help build their products. And then with digital natives, as they think about certainly building their native products, but then reducing their token cost for their software engineers as they are building their products up.
33 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000775467582