**Anjney Midha** (0:00)
Did you say hundreds of trillions?
**Jiaming Song** (0:01)
Yes.
**Anjney Midha** (0:02)
Let's be clear here. The world's largest open source model, Luma 3, was trained on 15 trillion tokens.
Dream Machine v0, your smallest model, is trained on hundreds of trillions of tokens.
**Jiaming Song** (0:16)
Yes.
I think currently there is still very active research on how to tokenize videos. And whereas in language modeling, I guess, the ways to tokenize things are relatively more mature. So I think we probably don't want to strictly use token count as the way of measuring the training data, but just use how big it is on the hard drive. And I can say with confidence that is nearly three out of 90 tubes larger.
**Derek** (0:46)
Welcome to the A16Z AI Podcast. I'm Derek. If you've been listening for a while, well, since we launched in April, you know that we like to highlight people who are building on the cutting edge of artificial intelligence. And this episode definitely checks that box. It features A16Z general partner, Anjney MIdha, in discussion with Jiaming Song, the chief scientist in a generative AI startup called Luma, and someone via his previous time as a Stanford PhD student and postdoc, and then at NVIDIA, who has played an outsized role in the development of generative video models.
You might have seen, and hopefully experimented with, Luma's recent Dream Machine launch, which generates 3D videos from text prompts. What you probably didn't realize, though, is the scale of data required to train such a model. At least hundreds of trillions of tokens, according to Song's best estimate of the raw video data Luma was working with. Throughout this chat, he explains how that's possible and how they train Dream Machine. He defines some key terminology in the 3D world, and he looks at where video and 3D models are headed. He also shows Anj some demos you might catch a reference to. Keep an eye on the a16z social media channels for some multimedia highlights.
But because it's a long and fascinating discussion, we're gonna get to it right now. As a reminder, please note that the content here is for informational purposes only, should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any a16z fund.
For more details, please see a16z.com/disclosures.
**Anjney Midha** (2:22)
Hey Jiaming, thanks for joining us.
**Jiaming Song** (2:24)
Hey, thanks Anj for the nice invite.
**Anjney Midha** (2:26)
Oh, I've been excited for this conversation for a while. I was thinking about on the way here, how to introduce you. And a story that came to mind was when I was at dinner sometime last year with a mutual friend of ours, Jim Fan from NVIDIA. And we started talking about the most interesting visual model research that was being done in the field. And of course your name came up. I remember either he said it or I said it, but I remember the quote being something like, behind every generated pixel today is a little bit of Jiaming.
You have had such a massive impact on the state of the art in diffusion models, in visual models. And I don't think most people realize just how many techniques that have allowed massive leaps in the quality of generative visual models of the last few years have your fingerprints all over them.
So why don't we just start by giving people a quick background on who you are and how we got here. Today, obviously, you're the chief scientist at Luma, but how did we get here?
**Jiaming Song** (3:29)
Yeah, thanks, Anjney, for the very nice introduction. And yeah, I'm Jiaming. So I'm currently the chief scientist of Luma AI.
Before Luma, I was a PhD student at Stanford, and I worked on machine learning with Stefano Erman. After that, I graduated around 2021, and then I went to continue to do a year of postdoc at Stefano's group, and then I went to NVIDIA to work on similar lines of topics on generative AI. Then around a year later, I felt that the opportunity is ripe for me to join Luma and to work on exciting things. But a little bit more on that detail on how I went into generative modeling or generative AI in general, I started working on similar lines of work during the undergrad. But at that time, machine learning wasn't really a big thing.
It was about 2014, and AlexNet just came out, and there was still a lot of resistance towards the deep learning paradigm, because it was so radical at that time.
51 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000661086342