**Anjney Midha** (0:00)
Thanks for listening to the a16z AI podcast. We have a fascinating and a lengthy discussion for you today, so we'll keep the introduction brief. If you're familiar with the world of generative AI models, you're likely familiar with LMArena, the leaderboard and competition space created and managed by a team at UC Berkeley. What began with the folks on language models has since expanded to cover vision models, coding models, and more. And very recently, the team behind LMArena announced they're starting a company to scale the project's reach and its impact. They want to amass a global community of AI users and use their collective experiences and ratings to make AI models more reliable and to help everyone find the right model for the right use case. So without further ado, here are LMArena founders Anastasios N. Angelopoulos, Wei-Lin Chiang, and Ion Stoica discussing the state and future of AI evaluation with a16z general partner Anjney Midha. They kick off the discussion discussing the importance of mass-scale real-time testing and evaluation right after these disclosures. As a reminder, please note that the content here is for informational purposes only, should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any a16z fund. For more details, please see a16z.com/disclosures.
**Anjney Midha** (1:26)
Sometimes, I get asked, what's the last exam that AI should take for humanity? And it seems like that's the wrong question to ask. We should be asking, what's the real-time exam you want your AIs to be taking before they get deployed every hour, every second of the day, especially as we start to get... I think one of the things that's emerging for me is that one of the arenas misunderstood partly because we're just early in AI.
And so while benchmarks like MMLU and the idea of these static exams were useful three years ago, the future is about real-time evaluation, real-time systems, real-time testing in the wild. Now, one thing that concerns a lot of people is the reliability of these systems. When we start going from chatbots that are good at, let's say, companionship and more consumer use cases to mission-critical systems, defense, healthcare, financial services, how will arena have to evolve as we go beyond companionship or web dev to those kinds of mission-critical use cases?
**Wei-Lin Chiang** (2:26)
I think that's one of the very reasons we wanted to create a company to support this project, to further scale the platform. So right now, we are at a million-month user now. What if we scale it to five, to ten, or even more, to capture even more diverse user base across different industries? And then in that case, we'll have ability to really zoom in into all these different areas that people really care about for critical mission task that it will be used to.
**Ion Stoica** (3:05)
You can imagine when we are going to scale, we can have micro-size for nuclear physicists, radiologists, and so forth, right? And these experts are going to come there to get the best answers to their, again, research questions.
**Anjney Midha** (3:22)
So that's interesting. Is there a future where, now that arena is becoming a company, you could see a scientific lab or a shipping company or a defense company deploy their own arena on their own infrastructure for their own users, on their own prompts?
**Anastasios N. Angelopoulos** (3:37)
Yeah, many people have asked us for this already.
**Anjney Midha** (3:40)
So these would be sort of private areas.
**Ion Stoica** (3:41)
Private evaluation.
**Anastasios N. Angelopoulos** (3:42)
And it's worth saying, I think when people have these mission critical industries in mind, they often are thinking about the factual nature of the responses and so on and so forth. But in reality, even in such industries, the majority of questions that people ask are subjective. Okay, so the mythology that in hard sciences or in mission critical industries, people just have like cut and dried questions and they just need like a retrieval and a lookup, that's completely false. That's the very reason why these models are useful. Because they allow you to sort of like interpolate between these weird questions and answer questions that are not fully specified and give responses that are sort of like geared to answer the question, but might not have a fully factual basis, right? And they might incorporate factual elements through RAD, let's say, but there's a subjective nature to the response. And that's a reality that everyone's going to have to live with. If these systems are going to be deployed in medicine and defense and so on and so forth, they're going to be deployed in places where the data is messy, because that's where they're useful. Okay, given that fact, how are you going to make sure that they're reliable? Well, you need something like Arena.
85 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000710577136