State of the Art: Training >70B LLMs on 10,000 H100 clusters artwork

State of the Art: Training >70B LLMs on 10,000 H100 clusters

Latent Space: The AI Engineer Podcast

June 25, 2024

It’s return guest season here at Latent Space!
Speakers: Shawn Pang, Josh Albrecht, Jonathan Frankle
**Shawn Pang** (0:00)
Welcome to the Latent Space Podcast, another super special edition.
Today, we have sort of like a two-header, Jon Frankle from Mosaic Databricks, or Databricks Mosaic, and Josh Albrecht from Imbue. Welcome.

**Josh Albrecht** (0:12)
Hey, glad to be here.

**Jonathan Frankle** (0:14)
Thank you for having us.

**Shawn Pang** (0:16)
Hey, so both of you are kind of past guests.
Jonathan, you were actually one of the most popular episodes from last year, talking about MPT 70B. Remember the days when we trained large models and there were 7B.

**Jonathan Frankle** (0:30)
Yeah, when reproducing Llama 1 7B was considered a huge accomplishment for the field. Those are the good old days. I miss that.

**Shawn Pang** (0:38)
That's the things have accelerated a lot. Actually, let's do a quick catch up and Josh, you can chime on in as well. So Databricks got acquired. I talked to you at-

**Jonathan Frankle** (0:45)
Mosaic got acquired, although-

**Shawn Pang** (0:47)
Sorry.

**Jonathan Frankle** (0:47)
Although sometimes it feels like Mosaic acquired Databricks because, you know, we're having a lot of fun being here, but you know, yeah.

**Shawn Pang** (0:52)
Yeah, I mean, you are chief scientist now of Databricks.

**Jonathan Frankle** (0:55)
Chief AI scientist. Careful with the title.
As much as I would love to understand how Spark works, I'm going to have to defer that to much smarter people than me.

**Shawn Pang** (1:03)
Got it. And I don't know about like what you would highlight so far as post acquisition, but the most recent news is that you guys released DBRX. Is that the thing that most people should be aware of?

**Jonathan Frankle** (1:13)
Actually, that's no longer the most recent news. Honestly, the most recent news, we announced this, but it was at our Data and AI Summit last week, so it was announced among like a hundred thousand other things, is that we finally released our text image model, which has been a year in the making through a collaboration directly with Shutterstock.
There was a lot of work put into finding a data set that we were comfortable with working on and trying to build a model that honestly, I felt like I could trust and that others might be able to trust to put out in the world. So that model was released last week. It's unfortunately just available via API due to the fact that the data is quite sensitive and quite valuable. It's Shutterstock's entire business in a lot of ways, but I'm still really excited that there's now a model that is trained on a data set where the provenance of every single image is known, and it's a damn good model, so I'm really proud of the team on that.

**Shawn Pang** (1:55)
Yeah, amazing. Josh, do you have any thoughts on image model questions?

**Josh Albrecht** (1:59)
That is not my area of expertise, but I was excited to see the release of it last week as well, and I'm very happy that you guys did a nice job on the data side of everything there, so that was cool to see.

**Shawn Pang** (2:09)
I think what's unusual is, I think Shutterstock's doing multiple deals in multiple labs.
So what is the Shutterstock model? Like, I guess, is this the house model for Shutterstock? Is this Databricks' version of the Shutterstock model? Like, what is this?

**Jonathan Frankle** (2:22)
The way that I would think about it is that Shutterstock is doing an amazing business in AI across the board. Their data set is kind of widely known to be the best stock photo data set in the world, the most comprehensive, the biggest. When you think about, like, what data set am I gonna train a multimodal model on, you call Shutterstock. And I at least, I've heard in the news, like OpenAI, Google, Meta, Apple have all called Shutterstock and made those deals. So a lot of models have had Shutterstock data incorporated into them, but this is the only model I know of so far where it was exclusively and specifically trained just on the vanilla Shutterstock data. There was nothing else mixed in. We didn't go and scrape the web and find other data or combined data sets or anything like that. And so this is in some sense, the house blend, but the other piece is that it's just a data set where the provenance of every image is known in public.
Where did the data come from? It is the Shutterstock collection. That's it, nothing less, nothing more. And certainly being at Databricks, if I've learned one thing, I've learned about enterprise customers and what they want out of AI. And one of the things they ask for most is just, what can you tell me about the data the model was trained on?

93 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000660208075