**Patrick O'Shaughnessy** (0:00)
This episode of Founders Field Guide is sponsored by Klaviyo. Want to deliver marketing moments that last a lifetime? Klaviyo is the ultimate marketing platform for e-commerce. With targeted segmentation, email automation, SMS marketing and more, Klaviyo helps you create your ideal customer experience. See why more than 50,000 brands, like Living Proof, Solo Stove and Nomad, trust Klaviyo to grow their business. Keep your customers coming back. Get a free trial at klaviyo.com/founders. That's klaviyo.com/founders. Stay tuned at the end of the episode where I talk to Klaviyo customer Nomad on their origin story and how they work with Klaviyo. This episode is also brought to you by Vanta. Does your startup need a SOC 2 report to close big deals? Or do you already have a SOC 2 report and want to make it easier to maintain? Vanta has built software that makes it easier to both get and renew your SOC 2 With Vanta's continuous monitoring solution, you avoid hosting auditors on site and taking hundreds of screenshots to prove that you are compliant, so you can focus on building your business. Vanta partners with audit firms who file your SOC 2 report directly inside of Vanta at a fraction of the normal cost.
Hundreds of companies, including more than 100 Y Combinator businesses, are leveraging Vanta's today to streamline compliance and focus on building their businesses.
Founder's Field Guide listeners can redeem a $1,000 off coupon at vanta.com forward slash Patrick. That's vanta.com forward slash Patrick.
Hello and welcome everyone. I'm Patrick O'Shaughnessy and this is Founders Field Guide. Founders Field Guide is a series of conversations with founders, CEOs and operators building great businesses.
I believe we are all builders in our own way and this series is dedicated to stories and lessons from builders of all types. You can find more episodes at investorfieldguide.com.
**SPEAKER_2** (1:46)
Patrick O'Shaughnessy is the CEO of O'Shaughnessy Asset Management. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of O'Shaughnessy Asset Management. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions.
Clients of O'Shaughnessy Asset Management may maintain positions in the securities discussed in this podcast.
**Patrick O'Shaughnessy** (2:11)
My guest today is Ali Ghodsi, founder and CEO of Databricks, a data analytics platform for data scientists and developers.
He's also the founder of Apache Spark, the open source project that Databricks is built on and is an accomplished researcher at UC Berkley's Computer Science Department. Our conversation ranges from the origins of distributed computing to modern data infrastructure, how companies can leverage their massive datasets and the transformation of Databricks through its phases of growth as a business. While technical, it's exactly the kind of conversation I like to have on this show. I hope you enjoy my great conversation with Ali Ghodsi. So Ali, I'd love to start our conversation at the end with what Databricks is today to level set for the audience exactly what you do, what your focus is and what the business does for customers. Could you just walk us through as we sit here at the end of 2020 what the company looks like and the service or problem it solves for customers?
**Ali Ghodsi** (3:04)
We're a seven-year-old company. We have about 1700 employees and we help enterprises take massive amounts of data and do machine learning, AI and data service on that data. Most enterprises, they've seen how Silicon Valley forward tech companies have used data in a really strategic way to disrupt industries. They want to do the same thing, but they don't have thousands of engineers that can help them build data platform custom for their use case.
We've built that and we enable them to do that.
**Patrick O'Shaughnessy** (3:34)
I would love to go all the way back and sort of tell the history of distributed computing because everybody will have heard the term big data. This was a really popular term, I don't know, five, seven years ago.
And I think that concept, that term, the fact that it was being talked about in normal business circles was the result of progress in the world of distributed compute and storage. I'd love you to rewind however far back you think is appropriate to go. Maybe it's back to the 2006 Yahoo days. Tell the modern history of distributing computing, what it means and why it's so interesting and important.
**Ali Ghodsi** (4:06)
I think what happened is that around 2000, we hit this wall, we call it Moore's wall, because they didn't figure out how to make computers faster. So everything started moving into these data centers.
New computer and it was a new data center. In these data centers where you had hundreds or thousands of machines, people started collecting more and more data. And the reason for this was multiple. One was the price of storage kept going down. So it became cheaper and cheaper to store all this massive amounts of data and no one wanted to throw it away. And they had heard that there were some poor tech companies like Google that had gotten a lot of value out of the data. So they wanted to do the same thing.
47 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000506860622