**Grace Gong** (0:00)
It's flexible.
Okay. We're live. Hi, Andrew. Welcome to Venture with Grace.
**Andrew Dai** (0:06)
Hi, Grace. Great to meet you. Thanks for having me on this.
**Grace Gong** (0:10)
Amazing. Before we start our first show, I want to give a quick shout out to our amazing sponsor. This episode is brought to you by Nebius, the ultimate cloud for AI innovators. Nebius provides AI infrastructure you can count on, combining reliability and speed with flexibility and engineering support unmatched by hyperscalers. AI leaders like Meta, Shopify, and Higgsfield are already partnered with Nebius to run their AI workflows. Plus, venture backed startup can save up to 150K on compute costs when they apply for access. Visit nebius.com or nebius.com/startup to learn more.
Okay, I want to start with your career journey. So you are the ultimate researcher as you help co-wrote the amazing paper that, oh, by the way, happy birthday. It's your birthday, I just noticed.
Oh my God, I love that. So actually, it's my mom's first day as well, but like, okay, we'll get there. But so I want to start with your journey from being a researcher at DeepMind for the past 13 years. What were some core lessons that you've learned early on in your career kind of shaped you into who you are today?
**Andrew Dai** (1:18)
That's a great question. So when I started at Google Brain, that was 12 years ago. Jeff Hinton and Elias Saskiafer were there as well. And one of the things that I do remember Jeff Hinton, well, like I was inspired by Jeff Hinton, was that a lot of the research he did and a lot of the ideas he had were inspired by how the human brain thinks. And basically that's the only example of intelligence we know of.
So it makes a lot of sense that we should use that to guide how we build our models, how we build AI. And that's turned out to be very fruitful. Like one of the theories we had when we wrote the paper on language model pre-training was that the human brain does when you listen to someone speak, you're actively predicting what they will say next.
Because then it kind of like reduces the cognitive load on your brain as well. So that's really being like a fundamental part of what we do. And then another thing I learned during my time at Brain and DeepMind is to not do research in isolation, to always do research with a view to having impact on real people, on the real world. And so I worked on products as well during my time at Brain, trying to get that research into products like Smart Reply and Smart Compose, which I worked on on YouTube as well. And then Google Health too, where we try to apply language models to Google Health back in the day.
**Grace Gong** (3:05)
For sure.
I guess like which sector do you feel like it will be most impacted by AI since you mentioned you work at Google Health, and obviously you work on Gemini and all these other frontier products. So how do you think about the sector that you want to start with since you guys are building a multimodal company?
**Andrew Dai** (3:28)
For the sector that we want to start with, we've seen that video understanding is quite a big area. There are so many applications within that. So that's one area we're looking into, and that requires quite a good deal of reasoning. So describing what happens in the video, and even in robotics, that's a critical use case because that can let you control the robot, whether it's a robot arm or robot gripper or humanoid robot. The first step of that involves understanding video.
So yeah, kind of this video understanding and maybe entertainment areas and video understanding robotics are our first areas that we're looking at.
**Grace Gong** (4:24)
For sure.
I want to start with, I guess one of the things you mentioned, it's like video is a really important component. Maybe we could talk about in the past, how does the previous multimodal companies are actually processing the video. So we have the founder of 12 labs on our podcast before. So essentially what he was saying is like in the past, when you look at Google or YouTube, we look at YouTube, it's like people have to translate the video into words or terms. So let's say I'm playing basketball and then the video have to convert into a text. And the text, when you're searching the transcription and you're like, hey, playing basketball or whatever, and then go back to the basketball area of the video. But 12 lab is doing essentially matching the video to the video. So if you are showcasing a video and then I guess it's just speed up the processing time. And I wonder, like, how do you, if you could like explain like Elorian to the normies, how would you kind of like go about it?
34 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000778342936