The Data Foundry for AI with Alexandr Wang from Scale artwork

The Data Foundry for AI with Alexandr Wang from Scale

No Priors: Artificial Intelligence | Technology | Startups

May 22, 2024

Alexandr Wang was 19 when he realized that gathering data will be crucial as AI becomes more prevalent, so he dropped out of MIT and started Scale AI.
Speakers: Elad Gil, Alexandr Wang, Sarah Guo
**Elad Gil** (0:05)
Hi, listeners, and welcome to No Priors. Today, I'm excited to welcome Alex Wang, who started Scale AI as a 19-year-old college dropout. Scale has since become a juggernaut in the AI industry.
Modern AI is powered by three pillars, compute, data, and algorithms. While research labs are working on algorithms, algorithms and AI chip companies are working on the compute pillar. Scale is the data foundry, serving almost every major LLM effort, including OpenAI, Meta, and Microsoft.
This is a really special episode for me, given Alex started Scale in my house in 2016, and the company has come so far. Alex, welcome. I'm so happy to be talking to you today.

**Alexandr Wang** (0:49)
Thanks for having me. Known you all for quite some time, so excited to be on the pod.

**Elad Gil** (0:53)
Why don't we start at the beginning just for a broader audience. Talk a little bit about the founding story of Scale.

**Alexandr Wang** (0:58)
Right before Scale, I was studying AI machine learning at MIT. This was the year when DeepMind came out with AlphaGo, where Google released TensorFlow. So maybe the beginning of the deep learning hype wave or hype cycle.
I remember I was at college, I was trying to use neural networks. I was trying to train image recognition neural networks. And the thing I realized very quickly is that these models were very much such as a product of their data. And I played this forward and thought through it. And these models or AI in general is the product of three fundamental pillars. There's the algorithms, the compute and the computation power that goes into them, and the data.
And at that time, it was clear there were companies working on the algorithms, labs like OpenAI or Google's Labs or a number of AI research efforts. There were, NVIDIA was already a very clear leader in building compute for these AI systems.
But there was nobody focused on data. And it was really clear that over the long arc of this technology, data was only going to become more and more important. And so in 2016, I dropped out of MIT, did YC, and really started Scale to solve the data pillar of the AI ecosystem and be the organization that was going to solve all the hard problems associated with how do you actually produce and create enough data to fuel this ecosystem. And really, this was the start of Scale as the data foundry for AI.

**Elad Gil** (2:32)
It's incredible foresight because you describe it as the beginning of the deep learning hype cycle. I don't think most people noticed that a hype cycle was yet going on.
And so I just distinctly remember you working through a number of early use cases, building this company in my house at the time, and discovering, I think far before anybody else noticed, that the AV companies were spending all of their money on data. Talk a little bit about how the business has evolved since then, because it's certainly not just that use case today.

**Alexandr Wang** (3:07)
AI is an interesting technology, because it is, at the core mathematical level, such a general purpose technology. It's basically functions that can approximate nearly any function, including intelligence.
And so it can be applied in a very wide breadth of use cases. And I think one of the challenges in building in AI over the past, we've been at it for eight years now, has really been what are the applications that are gaining traction and how do you build the right infrastructure to fuel those applications. So as an infrastructure provider, we provide the data foundry for all these AI applications. Our burden is to be thinking ahead as to where are the breakthrough use cases in AI going to be and how do we basically lay down the tracks before the sort of freight train of AI comes rolling through.
When we got started in 2016, this was the very beginning of the autonomous vehicle sort of cycle.
It was, I think, right when we were doing YC was when Cruz got acquired. It was sort of the beginning of the sort of the wave of autonomous driving being one of the key tech trends. And I think that we followed the early startup advice. You have to focus early on as a company. And so we built the very first data engine that supported sensor-fused data, so support a combination of 2D data plus 3D data, so lidars plus cameras that were built on onto the vehicles.
And then that very quickly became an industry standard across all the players, working folks like General Motors and Toyota and Stellantis and many others. In the first few years, the company were just focused on autonomous driving and a handful of other robotics use cases. But that was sort of the primetime AI use case.

33 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000656386075