The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality artwork

The Timeline for Realistic 4-D: Devi Parikh from Meta on Research Hurdles for Generative AI in Video and Multimodality

No Priors: Artificial Intelligence | Technology | Startups

July 20, 2023

Video dominates modern media consumption, but video creation is still expensive and difficult. AI-generated and edited video is a holy grail of democratized creative expression. This week on No Priors, Sarah Guo and Elad Gil sit down with Devi Parikh.
Speakers: Sarah Guo, Devi Parikh, Elad Gil
**Sarah Guo** (0:05)
Text prompts are democratizing creative expression, and the holy grail is AI generated and edited video.
Elad Gil and I sit down with Devi Parikh. She's a research director in generative AI at Meta, a leading researcher at multimodality in AI for visual, audio, and video. And she's an associate professor in the School of Interactive Computing at Georgia Tech. Recently, she worked on Make a Video 3D, which creates animations from text prompts. She's also a talented artist herself. Devi, welcome to No Priors.

**Devi Parikh** (0:35)
Thank you. Thank you for having me.

**Sarah Guo** (0:37)
Let's start with your background and how you got started in computer vision.
I've heard you say you choose projects based on what brings you joy. Is that how you got into AI research?

**Devi Parikh** (0:48)
Kind of, kind of, yeah. So my background is that I grew up in India, and then I moved to the US after high school, and I went to a small school called Rowan University in Southern New Jersey for my undergrad. And that is where I first got exposed to, what at the time was being called pattern recognition. We weren't even calling it machine learning, and got exposed to some research projects. There was a professor there who kind of showed some interest in me, thought I might have potential to contribute meaningfully to research projects.
And that's how I got exposed, and I really, really enjoyed what I was doing there. Decided to go to grad school, to Carnegie Mellon.
I knew I was enjoying it, but I wasn't sure if I wanted to do a PhD. So at first, I wanted to just kind of get a master's degree with a thesis, but I can do some research. But the year that I applied, that the ECE department at CMU decided that there wasn't going to be a master's track for thesis, that either you can just take courses or you go to a PhD. And so they kind of slotted me onto the PhD track, which I wasn't so sure of, but my advisor there was reasonably confident that I'm going to enjoy it, that I'm going to want to keep going. So yeah, that's how I got started in this space. At first, I was doing projects that didn't have a visual element to it.

**Sarah Guo** (2:02)
How did you pick a thesis project?

**Devi Parikh** (2:04)
So at first, I was working on projects that didn't have too much of a visual element to them. But when I got to CMU, my advisor's lab was working in image processing and computer vision. And I always thought that it was pretty cool that everybody gets to kind of look at the outputs of their algorithms and see what they're doing. But if it's kind of non-visual, then yeah, you see these metrics, but you don't really have a sense for what's happening, if it's working, if it's not.
And so that's how I got interested in computer vision and that then defined the topic of my thesis over the course of my PhD.

**Sarah Guo** (2:39)
So you have been working in machine learning long enough that as you said, it was called pattern recognition and you've worked across a bunch of different modalities. How has that changed your research path? Because things like diffusion models and GANs and large transformers, none of that existed when you were first starting. And I think you have managed to sort of translate or transition your interest in a way that keeps you on the cutting edge.
How does that happen?

**Devi Parikh** (3:09)
Yes, I think, I mean, you can always kind of look back and try and find patterns. Like when you're actually doing it, you don't necessarily have a grand strategy of anything in mind. But when I look back, I think one common theme that led to me transitioning across topics a little bit was that I was interested in seeing how we can get humans to interact with machines in more meaningful ways.
And so kind of even my transition from kind of non-visual to visual modalities in hindsight, I feel like was essentially that. I felt like you can't interact with these systems too much if it's sort of these abstract modalities that you're looking at. And then when I was working in computer vision, I wanted to find ways for humans to be able to interact with these systems more. So I started looking at kind of these attributes and adjectives of like, oh, something is funny or something is shiny. And using that as a mode of communication between humans and machines, both for humans to teach machines new concepts and for machines to be more interpretable in explaining why they're making the decisions that they're making. And that slowly led to the sort of more into natural language processing, where instead of these kind of just adjectives and attributes looking at more natural language as a way of interacting. So a lot of my work in visual question answering, where you're answering questions about images, image captioning was coming from there. And then over time, I sort of started thinking of, are there ways to go even deeper in this interaction? Are there ways where AI tools can enhance sort of creative expression for people, give them more tools for expressing themselves?

36 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000621746178