OpenAI’s Sora team thinks we’ve only seen the "GPT-1 of video models" artwork

OpenAI’s Sora team thinks we’ve only seen the "GPT-1 of video models"

No Priors: Artificial Intelligence | Technology | Startups

April 25, 2024

AI-generated videos are not just leveled-up image generators. But rather, they could be a big step forward on the path to AGI.
Speakers: Elad Gil, Tim Brooks, Aditya Ramesh, Bill Peebles, Sarah
**Elad Gil** (0:06)
Hi, listeners. Welcome to another episode of No Priors. Today, we're excited to be talking to the team behind Open AI's Sora, which is a new generative video model that can take a text prompt and return a clip that is high-definition, visually coherent, and up to a minute long.
Sora also raised the question of whether these large video models are world simulators, and applied the scalable transformers architecture to the video domain. We're here with the team behind it, Aditya Ramesh, Tim Brooks, and Bill Peebles. Welcome to No Priors, guys.

**Tim Brooks** (0:39)
Thanks so much for having us.

**Elad Gil** (0:40)
To start off, why don't we just ask each of you to introduce yourselves so our listeners know who we're talking to, Aditya, mind starting us off?

**Aditya Ramesh** (0:47)
Sure, I'm Aditya. I lead the Sora team together with Tim and Bill.

**Tim Brooks** (0:50)
Hi, I'm Tim. I also lead the Sora team.

**Bill Peebles** (0:53)
And Bill also lead the Sora team.

**Elad Gil** (0:55)
Simple enough, maybe we can just start with the Open AI mission is AGI, right, Greater Intelligence.
Is text to video on path to that mission? How did you end up working on this?

**Bill Peebles** (1:06)
Yeah, we absolutely believe models like Sora are really on the critical pathway to AGI. We think one sample that illustrates this kind of nicely is a scene with a bunch of people walking through Tokyo during the winter.
And in that scene, there's so much complexity. So you have a camera which is flying through the scene. There's lots of people which are interacting with one another they're talking, they're holding hands, there are people selling items at nearby stalls.
And we really think this sample illustrates how Sora is on a pathway towards being able to model extremely complex environments and worlds all within the weights of a neural network. And looking forward, in order to generate truly realistic video, you have to have learned some model of how people work, how they interact with others, how they think ultimately, and not only people, also animals and really any kind of object you want to model. And so looking forward, as we continue to scale up models like Sora, we think we're going to be able to build these world simulators, where essentially, anybody can interact with them. I, as a human, can have my own simulator running, and I can go and give a human and that simulator work to go do, and they can come back with it after they're done. And we think this is a pathway to AGI, which is just going to happen as we scale up Sora in the future.

**Elad Gil** (2:14)
It's been said that we're still far away, despite massive demand for a consumer product, like what is that on the road map? What do you have to work on before you have broader access to Sora?

**Tim Brooks** (2:27)
Yeah, so we really want to engage with people outside of Open AI and thinking about how Sora will impact the world, how it will be useful to people. And so we don't currently have immediate plans or even a timeline for creating a product. But what we are doing is we're giving access to Sora to a small group of artists as well as to Red Teamers to start learning about what impact Sora will have.
And so we're getting feedback from artists about how we can make it most useful as a tool for them, as well as feedback from Red Teamers about how we can make this safe, how we can introduce it to the public. And this is going to set our roadmap for our future research and inform if we do, in the future, end up coming up with a product or not, exactly what timelines that would have.

**Elad Gil** (3:11)
Aditya, can you tell us about some of the feedback you've gotten?

**Aditya Ramesh** (3:14)
Yeah, so we have given access to Sora to like a small handful of artists and creators, just to get early feedback.
In general, I think a big thing is just controllability. So right now, the model really only accepts text as input. And while that's useful, it's still pretty constraining in terms of being able to specify like precise descriptions of what you want. So we're thinking about like, you know, how to extend the capabilities of the model potentially in the future so that you can supply inputs other than just text.

**Sarah** (3:45)
Do you all have a favorite thing that you've seen artists or others use it for or a favorite video or something that you found really inspiring? I know that when it launched, a lot of people were really stricken by just how beautiful some of the images were, how striking, how you'd see the shadow of a cat in a pool of water, things like that. But I was just curious what you've seen sort of emerge as people, more and more people start using it.

28 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000653562932