The Mathematics of Training LLMs — with Quentin Anthony of Eleuther AI artwork

The Mathematics of Training LLMs — with Quentin Anthony of Eleuther AI

Latent Space: The AI Engineer Podcast

August 16, 2023

Invites are going out for AI Engineer Summit! In the meantime, we have just announced our first Actually Open AI event with Brev.dev and Langchain, Aug 26 in our SF HQ (we’ll record talks for those remote). See you soon (and join the Discord)!
Speakers: Alessio, Swyx, Quentin Anthony
**Alessio** (0:09)
Hey, everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO and resident at Decibel Partners, and I'm joined by my co-host, S.Wix, writer and editor of Latent Space.

**Swyx** (0:20)
Hey, today we have a very special guest, Quentin Anthony from EleutherAI. The context for this episode is that we've been looking to cover Transformers Math for a long time. And then one day in April, there's this blog post that comes out that literally is called Transformers Math 101 from Eleuther.
And this is one of the most authoritative posts that I've ever seen. And I think basically on this podcast, we're trying to give people an intuition around what are the rules of thumb that are important in thinking about AI and reasoning about AI. And I don't think there's anyone more credible than the people at Eleuther or the people training actual large language models, especially on limited resources. So welcome, Quentin.

**Quentin Anthony** (0:59)
Thank you. A little bit about myself is that I'm a PhD student at Ohio State University, starting my fifth year now, almost done.
I started with Eleuther during the GPT-NeoX 20B model. So they were getting started training that. They were having some problems scaling it.
As we'll talk about, I'm sure today a lot is that communication costs and synchronization and how do you scale up a model to hundreds of GPUs and make sure that things progress quickly is really difficult. That was really similar to my PhD work. So I jumped in and helped them on the 20B, getting that running smoothly. And then ever since then, just as new systems challenges arise and as they move to high performance computing systems and distributed systems, I just sort of kept finding myself falling into projects and helping out there. So I've been at Eleuther for a little bit now, had engineered there now and then finishing up my PhD. And then, well, who knows where I'll go next.

**Alessio** (1:47)
Awesome. What was the inspiration behind writing the article? Was it taking some of those learnings? Obviously Eleuther is one of the most open research places out there.
Is it just part of the DNA there or any fun stories there?

**Quentin Anthony** (2:00)
For the motivation for writing, you very frequently see in the DL training space these Twitter posts by, for example, Stas Bekman at Hugging Face. You'll see a Twitter post that's like, oh, we just found this magic number and everything is 20% faster. It's super excited but doesn't really understand what's going on.
And same thing for us. We very frequently find that a lot of people understand the theory or maybe the fundamentals of why AI training or inference works, but no one knows the nitty gritty details of how do you get inference to actually run correctly on your machine, split across two GPUs or something like that. So we sort of had all of these notes that we had accumulated and we're sort of sharing among engineers within Eleuther. And we thought, well, this would really help a lot of other people. It's not really maybe appropriate for like a paper, but for something like a blog post or technical report, this would actually maybe squeeze a lot of performance out of people's hardware they're already running on. So I guess there are a lot of projects in Eleuther that we're sort of trying to share notes with people in a way that typical institutions don't.
They sort of live within that institution, and then you go to a different institution and they do something very similar, but without the lessons of the previous. And it's because everyone's trying to do their own special sauce with their own stack. Whereas Eleuther, we don't really have that constraint and we can just share everything to everybody.

**Swyx** (3:14)
Yeah, this is a level of openness that basically very few people actually embrace. One, it's an extra effort to write things down, of course, but two, it is secret sauce. And so that not many people do it. And therefore, oftentimes the only way to learn this stuff is to actually work in one of the large model labs. And so you guys are doing a lot. The only other instance where I can think of where people actually open-source their process was Facebook's OPT.
What else is similar, like sort of trade knowledge, but not formal research knowledge?

**Quentin Anthony** (3:45)
I would say Bloom. So the Hugging Face Bloom project in big science and all of that, that was very open. I'd say it's the same caliber, if not more detailed than OPT.
Other than that, I think there was like a doc from Microsoft on like their Turing NLG. Their paper is pretty relaxed in that it did talk about some of those challenges. Other than like OPT and Bloom, I don't, and us, I don't, I can't think of any. It's a new thing.

47 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000624670587