**Alessio** (0:37)
Hey everyone, welcome to the Late in Space Podcast. This is Alessio, partner and CTO and resident at Decibel Partners, and I'm joined by my co-host, Swix, founder of Small AI.
**Swyx** (0:46)
Hey, and today we have a very special episode with Thomas Scialom. I don't know how to describe it. You've done so much work in a very short amount of time at Meta, but you were most notably leading Llama2, and now today we're also coordinating on the release of Llama3, so welcome.
**Thomas Scialom** (1:01)
Thanks for having me.
**Swyx** (1:02)
To be clear, obviously, the Llama3-405B, is that the official size number that we're going with? Or do we just say 400B?
**Thomas Scialom** (1:09)
For the text model only, yes. A bit of additional parameters for the multimodal version that will come later.
**Swyx** (1:17)
Awesome. Just to quickly go over your background, actually, we had a slightly similar past. I was also a quantitative trader, and it looks like you did five years in quant finance, working at Trading Timer in SockGen. Then you transitioned into natural language, getting your PhD at Sorbonne, working on Resset Tal as well, and then right after your PhD joining Meta.
**Thomas Scialom** (1:37)
It's exactly that, but basically, I think it's at the AlphaGo moment where I was doing some trading.
I say what I need to understand, what's the technology behind that, and I wanted to study machine learning. I did first some training, like six months degree, executive degree. At the end of which, I knew like what, eGboost at the time, and nothing about deep learning at all.
So, and most of the people around were like PhD people, and I was like, okay, PhD seems radical, deep learning seems radical, so I want to do a PhD in deep learning. That's where I joined, we have this PhD program in France within a company and academia. And so I did my PhD with Récital and Sorbonne University, on natural language generation and reinforcement learning. I guess it was a good topic. I was not like a visionary, it was very random. That's the company that offered me like this topic. And it was something like I started two weeks before birth.
**Swyx** (2:35)
Excellent timing. Yeah, we actually also just released our episode with Clementine Foguet, who also did her PhD with a company that kind of like a very similar format. I think, yeah, very underrated, very underrated, this sort of PhD with industry expertise, because you're also like publishing papers the whole time. I looked at your publishing history. You were doing like summarization work. You're doing factual consistency work. You released some benchmarks and he worked on language GANs before the Transformers took over.
**Thomas Scialom** (3:04)
We can come back to that later, but I should have papers have like 10, 50 citations. If I'm pretty sure that if I call them like.
RLHF without human in the loop, but like a discriminator which a synthetic human in the loop, I will have get much more citations today. Because like all the inspiration for this paper were from actually the original OpenAI paper of RLHF. But at Academia, we don't have the way to pay annotation online like that. So how to simulate it?
**Swyx** (3:38)
Yeah, a lot of these ideas are repeated like discriminator generator. We just call them different names now, like Verifier, whatever.
Well, I think your progress into NLP was like really strong. Because like the first thing you worked on at Meta was Bloom.
**Thomas Scialom** (3:50)
Yeah, actually, I started to work on that before joining Meta. I was not like one of the main contributors. But it was at the intersection of multilinguality, which was very important to me, large language modeling.
And that's why actually my first big project at Meta and the team I was working on was Galactica. And actually, an interesting step back from Bloom was like we did a lot of mistakes, but it was expression that's expected. And we learned a lot, but like trying to scale towards like multilinguality. In fact, we learned later that multilinguality almost emerged naturally with very, very few data, which was really surprising and not expected at all for us at the time.
**Swyx** (4:30)
I mean, my learning from that is just there's a natural harmony of language that is abstract from English.
When you learn English, you learn language, and that language just translates to other forms of languages, especially if they're the same family, right? Maybe we should get right into Llama2, spend a little bit of time there, and then we'll go into Llama3. What is the story of Llama2 from your point of view?
55 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000663106957