**Nathan Labenz** (0:00)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together, we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share my conversation with Anastasis Germanidis, CTO of RunwayML, a company that's become synonymous with AI video footage generation, and which, with their latest Gen 3 models, continues to lead the video generation market. Joining me as co-host for this episode is my good friend, Stephen Parker, creative director at Waymark and co-creator of the AI-generated film The Frost, which we've covered in a previous episode. As a power user of a great many generative art tools, including Runway, Stephen brings an extremely valuable perspective to this conversation. I came away from this discussion really super impressed with Anastasis. We threw a lot at him, and he gave very nuanced, thoughtful answers on a wide range of topics, including video generation models as world modelers, and the potential role of video understanding on the path to AGI. Emergent properties that Anastasis and the Runway team have observed as they've scaled up their models, including surprisingly accurate liquid simulations and improvements in 3D consistency, Runway's product development philosophy and culture of rapid iteration, including how they think about shipping upgraded models even if they might disrupt existing user workflows, and of course the challenge of scaling to meet explosive demand. We also touch on the creative possibilities unlocked by these tools, from pre-production storyboarding to generating footage for use in final productions, the importance of data quality over architectural complexity, the potential for the next generation of models to incorporate audio, the intriguing possibilities of interactive AI-generated environments, how all this technology might shape the future of entertainment and culture more broadly, and even how Anastasis thinks about competing in a game of scale with tech giants who have functionally unlimited resources.
In Runway's innovations, we can clearly see signs of things to come. Democratizing access to high-quality video creation means more stories from a wider range of voices, but also potentially major economic disruption for the iconic American film and television industry. As always, the pace of change is relentless. Since we recorded this episode, Runway has indeed launched API Access, and we at Weymark are super excited to be among the first wave of customers to try it out. If you're finding value in the show, we would love it if you'd take a moment to share it with friends or leave us a review on Apple Podcasts or Spotify. Of course, we welcome your feedback via our website, cognitiverevolution.ai or by DMing me on your favorite social network. Now, I hope you enjoy this illuminating discussion about the technology behind and the impact to come from AI video generation with Anastasis Germanidis, CTO of RunwayML. Anastasis Germanidis, CTO of RunwayML. Welcome to The Cognitive Revolution.
**Anastasis Germanidis** (3:11)
Great to be here.
**Nathan Labenz** (3:13)
I'm excited for this conversation and also excited to have my teammate and friend Stephen Parker, Creative Director at Waymark, along for this episode. We're going to get into generative AI for creative work and he is an expert in that with a film that we've covered in a previous episode, The Frost, now traveling the globe and making its debut in Singapore. I will be the least knowledgeable about what is going on here today, but excited to have this conversation. I thought we would start with a big picture discussion of video generation as world modeling and as a step on the perhaps critical path to AGI. I'd love to hear your perspective on the case that video generation really is something that we have to have on the way to an AGI destination.
**Anastasis Germanidis** (4:01)
So humans are visual beings, like being in the physical world is a fundamental aspect of being human, and you can formulate so many tasks humans do in the modality of video.
So the world models is like as those models basically learn through video data and large volumes of the video data, they gain powerful representations of the 3D world, they gain representations about a wide range of human activities, of different tasks that you can perform in the world, and so that knowledge can be leveraged to generate video, which is where we came from, but also for a variety of different other tasks. Video models will kind of power huge applications in robotics to build representations of the world. One framework I like to use is that every representation of the world, whether it's video, whether it's language, whether it's another presentation, it's ultimately a proxy for reality, but video itself has much less inductive biases than text, specifically the text that language models are trained on. Text captures a much smaller subset of everything that humans care about compared to video. So that's in a nutshell kind of the case for why video generation can lead to broadly useful general intelligence and how we're thinking about it.
45 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000672355737