**SPEAKER_1** (0:11)
Okay, we're here with Sarah Catanzaro from Amplify, welcome.
**Sarah Catanzaro** (0:14)
Thank you.
**SPEAKER_1** (0:15)
First time on the pod.
**Sarah Catanzaro** (0:15)
Great to be here. Took too long. I know, we've known each other for so long, yet never made an appearance.
**SPEAKER_1** (0:21)
It also made the transition from data to AI, I guess. I don't know if you were always as deep on AI, but obviously, there's a lot of simpatico.
**Sarah Catanzaro** (0:33)
Yeah. I've always actually kind of oscillated between data and AI. Sure. Arguably, I started my career in quote-unquote AI, it was just more symbolic systems back then. But as you said, I think they're so symbiotic, it's almost hard to divorce them. That's actually what brought me into data. I was like, I want to better understand what happens when I write a SQL query.
**SPEAKER_1** (0:56)
Yeah. Let's briefly touch on data, because I think obviously that's a lot of where you and I first met. DBT, Fivetran, that was so cool. I mean, how do you think about the end of the modern data stack?
**Sarah Catanzaro** (1:08)
Okay. So a lot of people look at the DBT, Fivetran merger and talk about the end of the modern data stack. And I think that is a fundamentally wrong take. Both of these companies were growing very healthily.
**SPEAKER_1** (1:27)
And you funded DBT?
**Sarah Catanzaro** (1:28)
We funded DBT. So both of the companies were actually beating their revenue targets. I think what you're more seeing is an IPO environment wherein companies are expected to have far more than 100 million revenue. And so...
**SPEAKER_1** (1:46)
What would you say the bar is now? 300?
**Sarah Catanzaro** (1:48)
No, above 600 Yeah, yeah.
**SPEAKER_1** (1:51)
And the combined company is 400?
**Sarah Catanzaro** (1:54)
I believe that they'll actually be close to 600 I don't have the exact number.
**SPEAKER_1** (1:58)
But they're clearly just getting ready for IPO.
**Sarah Catanzaro** (2:01)
So basically, the merger was a way to accelerate that path to liquidity. As you might remember...
**SPEAKER_1** (2:09)
And they were the presumptive winners in their categories anyway.
**Sarah Catanzaro** (2:11)
Exactly, exactly. I think one of the things that has actually pleasantly surprised me, and this speaks to, again, the symbiotic relationship between data and AI. Many of the big frontier labs are actually using both DBT and Fivetran. I recall talking to folks at Thinking Machines, within weeks of the company's formation, and DBT was already an important part of their stack. Certainly, training data sets need to be managed. We need insight into what users are doing on these platforms, and in fact, the way in which you would analyze interactions with an agent, or analyze interactions with an LLM is even more complicated.
While I think perhaps the demand for analytics engineers, the demand for data scientists didn't explode in the way that some people thought. Like, analytics engineers are not one third of personnel. That doesn't actually mean that the demand for the tools is not still very prevalent.
**SPEAKER_1** (3:14)
You got what you wanted. You wanted to democratize things. You got it.
**Sarah Catanzaro** (3:17)
Yeah. I guess we democratized things by perhaps reducing the need for the people. I don't know whether or not that is a good thing, but honestly, I do think that the fact that it is easier than ever from a tooling standpoint for people to make data-driven decisions is probably a step in the right direction. I've become actually convinced that, well, every company does need analytics engineers and does need data scientists. They probably don't need armies of them. Probably having a moderately sized data and analytics team is a good thing.
**SPEAKER_1** (3:55)
Yeah. You touched on an interesting thing. I wasn't planning to ask, but this is interesting. I come from the data field. The data was synonymous of analytics.
**Sarah Catanzaro** (4:03)
Yeah.
**SPEAKER_1** (4:04)
But you're now saying that the DBT Fivetran are being used for training data. Is there any notable differences in the workloads or the requirements?
**Sarah Catanzaro** (4:13)
Undoubtedly, there will be. I think one of the things that we saw with analytics that was surprising to some of the people in the data infrastructure space was that the workloads were actually quite predictable. They were quite predictable because many of them were actually not being generated by humans, but rather by deterministic systems. A lot of it was BI dashboards that are Tableau that is actually hitting your database or maybe not Tableau, but like Looker or Hax or something like that. I think with analyzing, curating, preparing datasets, it's a bit more ad hoc.
Undoubtedly, it will be less predictable. I don't know if that really changes the way that we approach developing data infrastructure. Some people are quite interested still in things like learned indexes, learned optimizers, and it's a bit easier to build a learned optimizer if you have more predictable workloads. It could change the way that we approach things like that.
21 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000748428004