Fei-Fei Li: Spatial Intelligence is the Next Frontier in AI artwork

Fei-Fei Li: Spatial Intelligence is the Next Frontier in AI

Y Combinator Startup Podcast

July 1, 2025

A fireside with Dr. Fei-Fei Li on June 16, 2025 at AI Startup School in San Francisco.Dr. Fei-Fei Li is often called the godmother of AI—and for good reason. Before the world had AI as we know it, she was helping build the foundation.
Speakers: Fei-Fei Li, Diana, Yashna, Carl, Annie
**Fei-Fei Li** (0:00)
My entire career is going after problems that are just so hard, bordering, delusional. To me, AGI will not be complete without spatial intelligence, and I want to solve that problem. I just love being an entrepreneur. Forget about what you have done in the past. Forget about what others think of you. Just hunker down and build. That is my comfort zone.

**Diana** (0:29)
So, I'm super excited here to have Dr. Fei-Fei Li. She has such a long career in AI. I'm sure a lot of you know her, right? Raise your hand.

**Fei-Fei Li** (0:43)
I know you too.

**Diana** (0:48)
She's been named the godmother of AI. One of the first projects that you created was ImageNet in 2009, 16 years ago.

**Fei-Fei Li** (1:02)
Oh my God. Don't remind me of that.

**Diana** (1:07)
Now, it has over 80,000 citations, and it really kicked off one of the legs of Stools for AI, which is the data problem. Tell us about how that project came about. It was pretty pioneering work back then. Yeah.

**Fei-Fei Li** (1:23)
Well, first of all, Diana and Gary and everybody, thanks for inviting me here. I'm so excited to be here because I feel like I'm just one of you. I'm also an entrepreneur right now. I just started a small company, so very excited to be here. ImageNet was, yeah, you're right. We actually conceived that almost 18 years ago. Time really flies. I was a first-year assistant professor at Princeton. Oh, wow. Hi. Hi, Tigers.
Yeah, and the world of AI and machine learning was so different at that time. There was very little data. Algorithms, at least in computer vision, did not work. There was no industry. As far as the public was concerned, the word AI doesn't exist. But there is still a group of us, starting from the founding fathers of AI, John McCarthy, and then we go through people like Jeff Hinton. I think we just had an AI dream. We really, really want to make machines to think and to work. And with that dream, my own personal dream was to make machines see. Because seeing is such a cornerstone of intelligence. Visual intelligence is not just perceiving, it's really understanding the world and do things in the world. So I was obsessed with the problem of making machines see. And as I was obsessively developing machine learning algorithms, at that time we did try neural network, but it didn't work. We pivoted to base net, to support vector machines, whatever it was. But one problem always haunted me. And it was the problem of generalization. If you work in machine learning, you have to respect that generalization is the core mathematical foundation or goal of machine learning. And in order to generalize these algorithms, these data, yet no one had data at that time in computer vision. And I was the first generation of grad student who was starting to dabble into data because I was the first generation of graduate student who saw the internet, the big internet of things. So fast forward around 2007-ish, my student and I decided that we have to take a bold bet. We have to bet that there needs to be a paradigm shift in machine learning, and that paradigm shift has to be led by data-driven methods. And there was no data, so we're like, okay, let's go to the internet, download a billion images, that's the highest number we can get on the internet, and then just create the world's, the entire world's visual taxonomy. And we use that to train and benchmark machine learning algorithm. And that was why ImageNet was conceived and came to life.

**Diana** (4:32)
And it took a while until there were algorithms that were promising. It wasn't until 2012 when AlexNet came out, and that was the second part of the equation with getting to AI, was getting the compute and throwing enough at it and algorithms. Tell us about what was that moment where you started to see, oh, you seeded it with data, and now the community started to figure more things out for AI.

**Fei-Fei Li** (4:59)
Right. So between 2009, we published this tiny little CVPR poster in 2009 to 2012, the AlexNet. There were three years that we really believe that data will drive AI, but we had very little signal in terms of if that was working. So we did a couple of things. One is we open-source. We believe from the get-go, we have to open-source this to the entire research community for everybody to work on this. The other thing we did is we created a challenge, because we want the whole world's smartest students and researchers to work on this problem. So that was what we call the ImageNet Challenge. So every year, we release a testing dataset. Well, the whole ImageNet is there for training, but we release testing, and then we invite everybody openly to participate. And then the first couple of years was really setting the baseline. You know, the performance was in the 30% error rate. It wasn't zero, or I mean, it wasn't completely random, but it wasn't that great. But the third year, 2012, I wrote this in a book that I published, but I still remember, it was around the end of summer, that we were taking all the results of ImageNet Challenge and running it on our servers. And I remember, it was late night. One day, I got a ping from my graduate student, I was home, and said, we got a result that really, really stood out, and you should take a look. And we looked into it. It was Convolutional Neural Network. It wasn't called Alex at that time, that team, that Jeff Hinton's team was called Supervision. It was a very clever play of the word super, as well as supervised learning. So Supervision. And we look at what Supervision did. It was an old algorithm. Convolutional Neural Network was published in the 1980s.

26 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000715301218