**Justin Johnson** (0:00)
I think the whole history of deep learning is, in some sense, the history of scaling up compute.
**Fei-Fei Li** (0:03)
When I graduated from grad school, I really thought the rest of my entire career would be towards solving that single problem, which is... A lot of AI as a field, as a discipline, is inspired by human intelligence. We thought we were the first people doing it. It turned out that was also simultaneously doing it.
**Justin Johnson** (0:26)
So Marvel, basically, one way of looking at it, it's a system, it's a generative model of 3D worlds, right? So you can input things like text or image or multiple images, and it will generate for you a 3D world that kind of matches those inputs. So while Marvel is simultaneously a world model that is building towards this vision of spatial intelligence, it was also very intentionally designed to be a thing that people could find useful today. And we're starting to see emerging use cases in gaming, in VFX, in film, where I think there's a lot of really interesting stuff that Marvel can do today as a product, and then also set a foundation for the grand world models that we want to build going into the future.
**Alessio** (1:12)
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swix, editor of Latent Space.
**Swix** (1:18)
And we are so excited to be in the studio with Fei-Fei and Justin of World Labs. Welcome.
**Fei-Fei Li** (1:24)
We're excited, too.
**Swix** (1:25)
I nearly said marble.
**Justin Johnson** (1:27)
Yeah, thanks for having us.
**Swix** (1:28)
I think there's a lot of interest in world models, and you've done a little bit of publicity around spatial intelligence and all that. I guess maybe one other part of the story that is a rare opportunity for you to tell is how you two came together to start building World Labs.
**Fei-Fei Li** (1:42)
That's very easy because Justin was my former student. So Justin came to my, you know, in my, the other hat I wear is a professor of computer science at Stanford. Justin joined my lab when? Which year?
**Justin Johnson** (1:57)
2012 Actually, the semester that I, the quarter that I joined your lab was the same quarter that AlexNet came out.
**Fei-Fei Li** (2:03)
Yeah. Yeah. So Justin is my first.
**Swix** (2:05)
Were you involved in the whole announcement drama? No.
**Justin Johnson** (2:08)
No, not at all. But I was sort of watching all the ImageNet excitement around AlexNet at that quarter.
**Fei-Fei Li** (2:14)
So he was one of my very best students.
And then he went on to have a very successful early career as a professor in Michigan, University of Michigan and Arbor in Meta. And then when we, I think around more than two years ago, for sure, I think both independently, both of us have been looking at the development of the large models and thinking about what's beyond language models. And this idea of building world models, spatial intelligence really was natural for us. So we started talking and decided that we should just put all the eggs in one basket and focus on solving this problem and started World Labs together.
**Justin Johnson** (3:02)
Yeah, pretty much. I mean, after seeing that ImageNet era during my PhD, I had the sense that the next decade of computer vision was going to be about getting AI out of the data center and out into the world. So a lot of my interests post-PhD shifted into 3D vision, a little bit more into computer graphics, more into generative modeling. And I thought I was drifting away from my advisor post-PhD, but then when we reunited a couple of years later, it turned out she was thinking of very similar things.
**Alessio** (3:31)
So if you think about AlexNet, the core pieces of it were obviously ImageNet, it was the move to GPUs and neural networks. How do you think about the AlexNet equivalent model for world models? In a way, it's an idea that has been out there, right? There's been, you know, Yann Le Guin is maybe like the most, the biggest proponent, most prominent of it. What have you seen in the last two years that you were like, hey, now's the time to do this? And what are maybe the things fundamentally that you want to build as far as data and kind of like maybe different types of algorithms or approaches to compute to make more models really come to life?
**Justin Johnson** (4:04)
Yeah, I think one is just there is a lot more data in compute generally available. I think the whole history of deep learning is in some sense the history of scaling up compute. And if you think about, you know, AlexNet required this jump from CPUs to GPUs, but even from AlexNet to today, we're getting about a thousand times more performance per card than we had in AlexNet days. And now it's common to train models not just on one GPU, but on hundreds or thousands or tens of thousands or even more. So the amount of compute that we can marshal today on a single model is, you know, about a million fold more than we could have even at the start of my Ph.D. So I think language was one of the really interesting things that started to work quite well the last couple of years. But as we think about moving towards visual data and spatial data and world data, you just need to process a lot more. And I think that's going to be a good way to soak up this new compute that's coming online more and more.
55 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000748427976