World Models and the Future of Spatial AI with Justin Johnson artwork

World Models and the Future of Spatial AI with Justin Johnson

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

September 1, 2026

In this episode, Justin Johnson, co-founder of World Labs, joins us to discuss world models and the emerging field of spatial AI.
Speakers: Sam Charrington, Justin Johnson

Topics: Technology, News, Tech News

**Sam Charrington** (0:01)
I want to send a big thanks to Blitzi for supporting the podcast and sponsoring this episode. Want to accelerate software development velocity by 5x? You need Blitzi, which brings autonomous software engineering to your enterprise. Blitzi starts by reverse engineering your code base, building a dynamic understanding of your entire application ecosystem. Your engineers simply declare intent, and once approved, Blitzi autonomously executes entire software epics, delivering validated end-to-end tested code. More than 80% of the work completed in a single run. Blitzi is not just generating code, it's developing software at the speed of compute. Experience Blitzi firsthand at blitzi.com/twiml.
That's blitzy.com/twiml.
The race to build more capable AI isn't just about making language models bigger. Increasingly, it's about world models, systems that understand space, predict how environments change, and act in the world around them. Justin Johnson is helping shape this shift. He co-founded World Labs with Fei-Fei Li and is an associate professor of computer science at the University of Michigan. I asked him why so many researchers think world models are AI's next frontier.

**Justin Johnson** (1:16)
There's a shared like low level belief among many researchers in the field that there's something that language models aren't doing, but that there's other kinds of models that we should be building, they do other kinds of things.
And that's something around understanding the world, generating worlds, simulating worlds, reconstructing worlds, planning actions through worlds. These are all capabilities that we want to build models to have. And why do we care about this is because we want to build systems that are not just stuck in a terminal or stuck as a virtual agent, right? You want to have visions of AI systems that are going to be robots that are out in the world acting in the world, or we want maybe want to build virtual worlds and live in those and simulate interesting things there. So all of these are capabilities that really don't feel like they're falling naturally out of the language modeling paradigm.

**Sam Charrington** (1:59)
I'm Sam Charrington and this is The TWIML AI Podcast. For over a decade, I've been exploring the ideas and innovation shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.
Over the past few years, one idea I think that has been espoused is essentially the idea that a world model is like an emergent property of either language models or diffusion. Like if you, you know, diffusion can, you know, illustrate some physical properties of the world or, you know, language models can tell stories that, you know, kind of seem like they know something about the world. And I guess there's a couple of questions emerging in my, you know, my question here. One is, like, maybe it's asking for a concrete definition, a concrete definition of a world model, because I think, you know, early on in those conversations, world model was really talking about, you know, these models having, like, foundational knowledge of the real world, like the world that we live in.
More recently, you know, the world model conversation, I feel like it shifted to talking about being able to create artificial worlds, but have them be self-consistent and navigable and, you know, properties like that. So, you know, A, kind of, I'd love to hear you, you know, elaborate on the relationship. And if you see kind of the same shift in the terminology, but also this idea of, like, you know, emergence and, you know, do we need new things or, like, if we throw enough, you know, data, compute, etc. at the models that we have, you know, does that get us there, or more likely, why won't that get us there?

**Justin Johnson** (4:08)
I think there's a lot of interesting questions to unpack there. One, I mean, the biggest one is just like, let's get it out of the way. There isn't a clear definition of world models that everyone in the field agrees on. And I think that's causing part of the confusion, right? Like, there is not, like, a thing where we can say a model that has X property or produces X kind of input and produces X kind of output, like, is a world model definitionally. I think we don't have that crisp definition as a field of what we mean, which leads to the confusion.
But to your point, I think there's a couple variants of this that feel like they're, like, I think there is some notion of, like, implicit world knowledge that you mentioned, that there are other kinds of models that produce certain kinds of inputs and outputs. Maybe if a model that is producing text, if it produces the right kind of text, the only way it could have produced this kind of text answer is because it knows something about the real world or because it's modeling some kind of implicit world internally in its neural network weights. Or similarly for video models, right? Like, if I am able to generate a video that is super photorealistic and has detail and like has physics and has water running in very particular ways, like maybe implicitly the model must have been modeling something about a real world in order to generate an output of that kind.

60 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID