**Dylan Fox** (0:00)
The amount of weekly conversations that Assembly handles through our APIs every week is up over 800 percent over the last three years. On a peak week, there'll be something like over 120 million voice conversations going through our platform, over two million hours of voice, which is a little over 4x the amount of daily volume that's going to YouTube. So they're getting more accurate, they're getting faster, more controllable, more capabilities. The newest voice models that we just launched, you can give them context about the environment they're operating in. If you drive through a McDonald's and you order with your voice, that voice AI system has no clue that it's taking a McDonald's order. But with our models, you can give them context on, hey, you're taking a McDonald's order, you want to focus on the person that's ordering, ignore the kids that are screaming in the background. So we're seeing this huge inflection now, like you have enterprises, small businesses that can build with API infrastructure now. Our TAM has just increased by 100x, because we're not just selling to engineering teams within product companies, it's now like anyone.
**Molly O'Shea** (1:06)
Dylan Fox, welcome to Sourcery.
**Dylan Fox** (1:07)
Yeah, thanks for having me here.
**Molly O'Shea** (1:09)
So we're going to have a very fun conversation on voice. We were just gabbing off-camera about how much I love voice, and how I've used it to start this podcast and do all these interviews. Specifically, use the transcripts that I capture from recordings into different kinds of formats. I started it for investing, and so I would take that, I would repackage it, I would make open-source expert calls and open-source memos, and all this kind of stuff just from capturing that data. Now, you have built the infrastructure for how that even works. So I want to get into that. I want to get into everything from the business model, how you've built out the products. You've been doing this for a very long time, and there's a lot of new entrants to the category, so it'd be great to learn what is BS and what's not, as well as areas that you're most excited about. But before we start, I guess it would be great to kind of capture and illustrate just how big AssemblyAI has gotten, how much data you train on.
**Dylan Fox** (2:08)
Yeah, so the demand for voice applications and the applications that companies are building on our infrastructure is, over the last two years in particular, really started to explode. One stat I was just looking at before I came over here was like, the amount of weekly conversations that Assembly handles through our APIs every week is up over 800 percent over the last three years. Oh my God. So now, on a given week, on a peak week, there'll be something like over 120 million conversations, voice conversations going through our platform, over two million hours of voice, which as of December of this past year, is 4X or a little over 4X the amount of daily volume that's going to YouTube. Which when I found that was like, wait, I need to double check this because that can't be right. But yeah, the volume is huge. There's almost 100 million API calls a day coming against our API, about a million developers, a little over a million developers on the platform now.
And 40 percent of those developers signed up to the API last year. And so while we started this a long time ago, it's really over the last like two, three years where voice and edits accelerating and inflecting voice is becoming just a core part of software, and increasingly hardware too. And so we're seeing that through our platform because we're the infrastructure really under like under all of it.
**Molly O'Shea** (3:34)
So if you look back over time, you guys went through YCE. Yeah.
Did you think that this would be the reason why voice would take off and people talking to their phones and talking to their computers and that kind of thing? What did you think?
**Dylan Fox** (3:48)
Yeah. I mean, I did, which is why I've spent so much time on this because we were actually the very first AI batch in YCE when we went through YCE. And so it was Daniel Gross, who if you know Daniel now at Meta, he started the AI batch at YCE back when we went through in 2017 And it was me and like five other companies. And we got $100,000 in GPU credits. That was like our perk, which seems like cute, right?
It's like, wow, that's nothing. But back then it was like, wow, $100,000 of NVIDIA K80 usage. This is amazing. But the AI ecosystem back then was just in its infancy. I was going to the very first TensorFlow meetups just to put it into perspective. No one is really using AI in production yet. But for me personally, I had got an Amazon Echo and was really into the voice interfaces and talking to this hardware that worked well. Because my experience with voice prior had been, everything was terrible.
37 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000779257340