**Arthur Mensch** (0:00)
If you look at the reason why we made all of this progress, most of it is explainable by the free flow of information. You had academic labs, you had very big industry-backed labs communicating all the time about the results and building on top of each other results. And all of a sudden in 2020, with GPT-3, this tide reversed and companies started to be more opaque about what they were doing because they realized there was actually a very big market. That's something that I, as a researcher, and all of the people that joined us as well, deeply regretted because we think that we're definitely not at the end of the story.
We believe that it's still the case that we should be allowing the community to take the models and make it their own.
**Anjney Midha** (0:41)
Hi, you're listening to the a16z AI Podcast. As we kick off a long holiday weekend, at least here in the United States, we want to share a couple of episodes from the main and self-titled a16z Podcast archive. All of these are much more recent than our last archive episode with GPT-3. Rather, we're stitching together two episodes featuring a16z general partner, Anjney Midha, interviewing a couple of guests working very much at the state of the art. The first episode from December, 2023 features Arthur Mensch, the co-founder and CEO of Mistral, arguably the world's foremost provider of open large language models. Anjney and Arthur discuss the importance of open models for advancing AI research and innovation, as well as Mistral's then new Mixtral 7b model, which applied a mixture of experts' approach to deliver very high performance.
And then, at about the 34-minute mark, you will hear Anjney in discussion with Stanford Professor Stefano Ermon. In this discussion from February, 2024, on the heels of OpenAI showing off its Sora video model, they discuss the state of generative AI video creation. Enjoy.
As a reminder, please note that the content here is for informational purposes only, should not be taken as legal, business, tax or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any a16z fund. For more details, please see a16z.com/disclosures.
**Arthur Mensch** (2:17)
If you look at the reason why we made all of this progress, most of it is explainable by the free flow of information. You had academic labs, you had very big industry-backed labs communicating all the time about the results and building on top of each other results.
And all of a sudden in 2020 with GPT-free, this tide reversed and companies started to be more opaque about what they were doing because they realized there was actually a very big market. That's something that I as a researcher and all of the people that joined us as well, deeply regretted because we think that we're definitely not at the end of the story. We believe that it's still the case that we should be allowing the community to take the models and make it their own.
**Anjney Midha** (2:58)
All right, why don't we start with the founding team story. We flashback to a few years ago, labs are building foundation models and the consensus across the research community was that the size of these models was what mattered most. How many million or billion parameters went into the model seemed to be the primary debate that people were having.
But you had a hunch that the role of data mattered more. Could you just give us the backstory on the chinchilla paper you co-wrote? What were the key takeaways in the paper and how was it received?
**Arthur Mensch** (3:30)
Yeah, so I guess the backstory is that in 2019, 2020, people were relying a lot on the paper called scaling laws for large language models. That was advocating for basically scaling infinitely the size of models and keeping a number of data points rather fixed. So just saying that if you had four times the amount of compute, you should be mostly multiplying by 3.5 your model size and then maybe by 1.2 your data size. And so a lot of work was actually done on top of that. So in particular, DeepMind, when I joined a project called Gopher and there was a misconception there, there was also a misconception on GPT-3. And basically in 2021, every paper made this mistake.
And at the end of 2021, we started to realize there was some issues when scaling up.
And as it turns out, we turned back to the mathematical paper that was actually talking about scaling those and it was a bit hard to understand. And we figured out that actually, if you thought about it a bit more in a theoretical perspective, and if we looked at the empirical evidence we had, it didn't really make sense to actually grow the model size faster than the data size.
61 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000656646281