In the Arena: How LMSys changed LLM Benchmarking Forever artwork

In the Arena: How LMSys changed LLM Benchmarking Forever

Latent Space: The AI Engineer Podcast

November 1, 2024

Apologies for lower audio quality; we lost recordings and had to use backup tracks.
Speakers: Alessio, Swyx, Anastasios Angelopoulos, Wei-Lin Chiang
**Alessio** (0:04)
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner in C2N Residence and Decibel Partners, and I'm Jordan, my micro-hostess, Wix, founder of Small AI.

**Swyx** (0:14)
Hey, and today we're very happy and excited to welcome Anastasios and Wei-Lin from LMSys. Welcome, guys.

**Anastasios Angelopoulos** (0:21)
Hey, how's it going? Nice to see you. Thanks for having us.

**Swyx** (0:24)
Anastasios, I actually saw you, I think, at last year's NeurIPS. You were presenting a paper which I don't really super understand, but it was some theory paper about how your method was very dominating over other search methods. I don't remember what it was, but I remember that you were a very confident speaker.

**Anastasios Angelopoulos** (0:40)
Oh, I totally remember you. Didn't ever connect that, but yes, that's definitely true. Yeah, nice to see you again.

**Swyx** (0:46)
Yeah, I was frantically looking for the name of your paper and I couldn't find it. Basically, I had to cut it because I didn't understand it.

**Anastasios Angelopoulos** (0:51)
Was this Conformal BID Control or was this the Home Control? Blast from the past, man.

**Swyx** (0:57)
Blast from the past. It's always interesting how in Europe and all these academic conferences are sort of six months behind what people are actually doing. But Conformal Risk Control, I would recommend people check it out. I have the recording, I just never published it just because I was like, I don't understand this enough to explain it.

**Anastasios Angelopoulos** (1:14)
People won't be interested, it's all good.

**Swyx** (1:16)
But ELO scores, ELO scores are very easy to understand. You guys are responsible for the biggest revolution in language model benchmarking in the last few years. Maybe you guys want to introduce yourselves and maybe tell a little bit of the brief history of LMSys.

**Wei-Lin Chiang** (1:31)
Hey, I'm Wei-Lin. I'm a fifth-year PhD student at UC Berkeley, working on Chatbot Arena these days, doing crowdsourcing, AI benchmarking.

**Anastasios Angelopoulos** (1:42)
I'm Anastasios. I'm a sixth-year PhD student here at Berkeley. I did most of my PhD on theoretical statistics and foundations of model evaluation and testing. And now I'm working 150% on this Chatbot Arena stuff. It's great.

**Alessio** (2:00)
And what was the origin of it? How did you come up with the idea? How did you get people to buy in? And then maybe what were one or two of the Pevoto moments early on that kind of made it the standard for these things?

**Wei-Lin Chiang** (2:11)
Yeah, yeah. Chatbot Arena project was started last year in April, May, around that. Before that, we were basically experimenting in the lab how to fine-tune a Chatbot open source based on the Lama 1 model may have released. At that time, Lama 1 was like a base model, and people didn't really know how to fine-tune it. So we were doing some explorations. We were inspired by Stanford's AlphaCop project. So we basically, yeah, grow a data set from the Internet, which is called SharedG2BT data set, which is like a dialogue data set between user and ChatG2BT conversation. And it turns out to be like pretty high quality data, dialogue data. So we fine-tuned it, and then we trended and released a model called VKUNIA. And people were very excited about it, because it kind of like demonstrate open-way model can reach this conversation capability similar to ChatG2BT. And then we basically released the model with, and also built a demo website for the model. That's people were very excited about it. But during the open, the biggest challenge to us at the time was like, how do we even evaluate it? How do we even argue this model we trend is better than others? And what's the gap between this open-source model that other proprietor offering? At that time it was like GPT-4 was just announced. It's like Cloud One. What's the difference between them? And then after that, every week, there's a new model being fine-tuned, released. So even until still not right. And then we have that demo website, Orbit Cuneo, and then we thought like, okay, maybe we can add a few more open model as well, like API model as well. And then we quickly realized that people need a tool to compare between different models. So we have a side-by-side UI implemented on the website to that people choose to compare. And we quickly realized that maybe we can do something like a battle on top of these OMS, like just anonymize it, anonymize the identity, and that people vote which one is better. So the community decide which one is better, not us arguing, our model is better or what. And that turns out to be like people are very excited about this idea.

35 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000675358435