The Future is Small Models, with Matei Zaharia, CTO of Databricks artwork

The Future is Small Models, with Matei Zaharia, CTO of Databricks

No Priors: Artificial Intelligence | Technology | Startups

April 6, 2023

If you have 30 dollars, a few hours, and one server, then you are ready to create a ChatGPT-like model that can do what’s known as instruction-following. Databricks’ latest launch, Dolly, foreshadows a potential move in the industry toward smaller and more accessible but extremely capable AIs.
Speakers: Matei Zaharia, Sarah Guo, Elad Gil
**Matei Zaharia** (0:07)
So we really wanted to see whether it's possible to democratize this and to let people build their own models with their own data without sending it to some centralized provider that's trying to learn from everyone's data and control their destiny in this space.

**Sarah Guo** (0:26)
This is the No Priors Podcast. I'm Sarah Guo.

**Elad Gil** (0:29)
I'm Elad Gil.

**Sarah Guo** (0:30)
We invest in, advise, and help start technology companies.

**Elad Gil** (0:33)
In this podcast, we're talking with the leading founders and researchers in AI about the biggest questions.

**Sarah Guo** (0:45)
If you have $30 a few hours in one server, then you're ready to create a chat GPT-like model that can do what's known as instruction following. The latest launch, Dolly, from Databricks, which is available in open source, foreshadows a potential move in the industry towards smaller and more accessible, but extremely capable AIs.
Matei Zaharia, co-founder and chief technologist at Databricks, is here to tell us all about Dolly. We'll talk about how big data sets actually need to be, why manual annotations becoming less necessary to train some models, and how he went from a Berkeley PhD student with a little project you may have heard of called Spark to the founder of a company that's now critical data infrastructure, that's increasingly moving into AI. Welcome to the podcast, Matei.

**Matei Zaharia** (1:22)
Thanks a lot. Excited to be here.

**Sarah Guo** (1:23)
Can you start by telling us a little bit about the origins of Databricks and how it led you to where you are today?

**Matei Zaharia** (1:28)
Sure, yeah. So Databricks started from a group of seven researchers at UC Berkeley back in 2013
And we were really excited about democratizing basically the use of large data sets and of machine learning. So we had seen the web companies at the time were very successful with these things, but most other companies, most other organizations, things like scientific labs and so on, weren't. And we were really excited to look at making it easier to do computation on large amounts of data and also to do machine learning at scale with the latest algorithms. So we had started doing our research. We worked with some of the web companies. We also started open source projects like most notably Apache Spark, which was essentially, the first version of it was my PhD thesis.
And we had seen a lot of interest in these. And we thought it would be great to start a company to really reach enterprises and make this type of thing much better and actually allow other companies to use this stuff.

**Sarah Guo** (2:27)
Can you just give us a sense of what Databricks looks like today from a scale and product suite perspective?

**Matei Zaharia** (2:33)
Sure, yeah. So Databricks offers a pretty comprehensive data and ML platform in the cloud. It runs on top of the three major cloud providers, Amazon, Microsoft and Google.
And it includes support for data engineering, data warehousing, machine learning. And most interestingly, all this is integrated into one product. So for example, you can have one definition of your business metric that you use in your BI dashboards, and the same exact definition is used as a feature in machine learning. And you don't have this drift or copying data, and you can just kind of go back and forth between these worlds.
The company has about 6,000 employees now, and last year we said that we crossed a billion dollars in ARR, and we're continuing to go. It's a consumption-based cloud model where, you know, customers that are successful can go over time and begin new use cases and so on.

**Sarah Guo** (3:25)
Did you think the opportunity was as big as it has been when you started the company?

**Matei Zaharia** (3:29)
Yeah, well, we definitely didn't, you know, anticipate necessarily to go to this size, right? A lot of things can go wrong, but we were excited about the confluence of a few trends. So first of all, you know, it's so easy to collect large amounts of data and people are doing it automatically in many industries.
And second, cloud computing makes it possible to scale up very quickly, do experiments, scale down and so on, which enables more companies to work with this kind of thing. And then the third one was machine learning. So we thought, you know, these are powerful trends. And the exciting thing for us as a company is we didn't invent cloud computing. We didn't necessarily invent big data or anything, but we were able to start at a point in time when many companies were thinking to move into this space and just provide a great platform for that. And there's this migration already happening.

34 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000607682628