Fei Fei Li: The Race to Build World Models For AI artwork

Fei Fei Li: The Race to Build World Models For AI

The a16z Show

September 4, 2026

World Labs co-founders Fei-Fei Li, Justin Johnson, and Ben Mildenhall join a16z General Partner Martin Casado to discuss Atlas, their latest world model, and what it reveals about the pursuit of spatial intelligence.
Speakers: Fei-Fei Li, Justin Johnson, Ben Mildenhall, Martin Casado

Topics: Technology, Business, Entrepreneurship

**Fei-Fei Li** (0:00)
On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded, that is the very hard step that Atlas has taken.

**Justin Johnson** (0:10)
We know LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction.

**Ben Mildenhall** (0:16)
This is the real place where AI can actually unlock a ton of value for people and their process. We're saying like 50, 100x reduction.

**Justin Johnson** (0:23)
There was a famous shot in the first Matrix movie where Neo is falling down. Exactly. They had hundreds of cameras doing that angle on a green screen. On Atlas, we can do this with just three cameras. No studio capture, no green screen, no expensive calibration.

**Fei-Fei Li** (0:36)
No one has ever seen this result.

**Martin Casado** (0:39)
When you set out to do this, did you know it was going to work?

**Justin Johnson** (0:41)
I was pretty sure. Each time we made the model bigger and each time we trained it for longer, it got significantly better.

**Martin Casado** (0:46)
Does that mean we're going to get 4D video? Like go walk around?

**SPEAKER_5** (0:50)
Language models are built around predicting the next token. What happens when a model instead learns to predict the next view of the world? In this episode, Martin Casado sits down with World Labs co-founders Fei-Fei Li, Justin Johnson and Ben Mildenhall to discuss Atlas and the broader challenge of building AI that can reason about physical space. They explain how Atlas combines generation and 3D reconstruction, why the team chose new view prediction as its underlying primitive, and what they learned trying to scale an approach that hadn't been tested before. They also get into what's still missing, including richer dynamics and interaction, and how world models could eventually connect simulation with robotics and planning.
Underlying it all is a bigger hypothesis. Could new view prediction play a similar role for spatial intelligence that next token prediction has played for language?

**Martin Casado** (1:44)
It's a big day yesterday. You launched a new frontier model, which has got an amazing reception, which is still coming in. I think maybe a good way to structure this conversation, let's just talk about exactly what that was, and then we'll go back to history and work our way back up. Maybe, Justin, you want to talk about what was launched yesterday, why it's significant.

**Justin Johnson** (1:58)
Yeah. Atlas is our new next generation world model. It has three basic things. It can generate, reconstruct and simulate the world. Within that, there's a couple of different major capabilities. It has really good camera condition generation. You can input an image together with the camera trajectory and steer the model and have it generate video frames from any perspective you want. It's really good at sparse 3D reconstruction. You can input one or multiple up to a hundred frames that are views of the real world and use those to reconstruct the real world. That reconstruction can take the case either of novel, a video flying through the space or an explicit 3D reconstruction of the space. Then finally, it can be used for simulation. For this, we show off these awesome bullet time videos, which got a lot of attention online, and then also robotic simulation.

**Martin Casado** (2:37)
What's a bullet time video?

**Justin Johnson** (2:39)
A bullet time video. This comes from The Matrix. There was a famous shot in the first Matrix movie where Neo was like, Oh, they're like, what?
Then you remember in that famous shot, he's falling down, it's in slow motion, and the camera flies all the way around. The way that they did that shot is they had a ring of hundreds of cameras. Then he fell over in the studio, they had hundreds of cameras viewing that angle on a green screen, and then they used those hundreds and hundreds of cameras to make that famous shot in Matrix. But now with Atlas, we can do this with three cameras. No studio capture, no green screen, no expensive calibration. We can literally stick three iPhones onto iPods, use these to take a video of something happening, like someone shooting a basket, someone dropping a strawberry into a bowl of milk. And then from those three iPhone videos, we can then reframe the shot and imagine like freeze time, have the camera fly in as the milk is splashing up and get these amazing frozen time views. And we can do this with just a couple of cameras.

**Martin Casado** (3:30)
What is the simplest description of what Atlas does, what goes in and what comes out?

41 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID