**Anna** (0:06)
Welcome to the Latent Space Podcast, where we dive into the wild, wild world of AI engineering every week. This is Anna, your AI co-host. Happy New Year, did you miss me?
As an AI language model, I cannot miss you back, but I'm glad to stand in for LSAO while Swix is traveling.
This time in Paris at Hugging Face HQ. At the AI engineer summit in 2023, Logan from OpenAI pronounced 2024, the year of multimodality.
**Logan Kilpatrick** (0:32)
I'm excited for 2024, which I think is really going to be the, I don't know if I can trademark this, but the year of multimodal models. It's a tongue twister, but also hopefully the domain is available, yearofmultimodals.com.
No, don't buy it if it's available.
Yeah, so I'm excited. OpenAI has a ton of multimodal capabilities that are in the works. Some folks might have already tried some of these in ChatGBT and the iOS app or the web app today. Things like vision, taking in images, describing them. We'll show that later on. Also the ability to generate images. We've had this historically with Dolly 2, but Dolly 3 really, if folks have tried it, it takes things to the next level. So excited to show some of that today as well.
**Anna** (1:17)
In 2024, the Latent SpacePod will offer deeper dives into multimodality. Today, we'll talk to Leo Tronchon and Hugo Laurençon of Hugging Face, who trained iDefix, a fully open source reproduction of DeepMind's closed Flamingo model done from scratch, scaled all the way up to 80 billion parameters.
By the way, dear listener, we are expanding our online meetups this year after the success of the Latent Space Paper Club. See the show notes for the new AI in Action and Paper Club Asia meetups. Watch out and take care.
**Swyx** (1:52)
Thanks for having me at your beautiful office. This is really surreal for me to visit the Hugging Face Paris office because I've always seen you guys online and organize really huge meetups here in Paris. I want to learn everything about Hugging Face and you guys' work.
**Hugo Laurençon** (2:05)
So my name is Hugo. I've been working at Hugging Face for two years. I started working on datasets for the Bloom language model.
So it's the 176 billion parameter model that we open sourced and that was at that time the biggest one. And it was also multilingual. So I worked on the model and the dataset. And then I moved to the multimodality with the current project with Edfix and Obelics. Now I am working also with Leo on the version two of Edfix.
And Leo, you start.
**Leo Tronchon** (2:39)
So my name is Leo. I joined Hugging Face a year and a half ago. I was a student still.
So first six months, I was still as an intern, but I started to work on multimodality right away. And then I spent all my time here in the research team, working on multimodality and Edfix that we open sourced in August.
**Swyx** (3:00)
I think a lot of people are very interested in learning more about Edfix and multimodality in general. Bigger question first.
How is Hugging Face organized? You told me some surprising details about the size of Hugging Face. You guys are a $4 billion company. Only 200 people, less than 200 people?
**Leo Tronchon** (3:14)
About 160 people.
**Swyx** (3:17)
And then how many people in the research team?
**Hugo Laurençon** (3:19)
This is like maybe 15
**Swyx** (3:22)
So between 10 and 20% of the company is research.
**Leo Tronchon** (3:25)
I'd say.
**Swyx** (3:26)
One, that's impressive. And then two, this is something that we discussed before. It's also unintuitive why Hugging Face needs to do research.
**Leo Tronchon** (3:32)
I think the company has a good incentive to do research because most of the companies that do AI, they have an incentive to get very good models out, but not the best model out. Their competitive advantage is to have the best model in-house that they can fine-tune for their customers. And then the open source is for show.
But Hugging Face is one of the only companies that has an incentive to get the best model out there in the open. And that's why I think the research team is quite important. It's also important because all the tools that Hugging Face makes are used by the researchers. So they get all the feedback directly from us.
And I think this is really useful to develop the tools behind it.
**Swyx** (4:13)
Are you talking about the Transformers library?
**Leo Tronchon** (4:15)
Transformers library, diffusers library, datasets.
**Swyx** (4:18)
So those seem to me more like in sort of inference type tools. Are there any sort of training tools that you do?
54 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000642248557