Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith artwork

Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith

Latent Space: The AI Engineer Podcast

January 8, 2026

Happy New Year! You may have noticed that in 2025 we had moved toward YouTube as our primary podcasting platform.
Speakers: Micah-Hill Smith, swyx, George Cameron
**Micah-Hill Smith** (0:06)
This is kind of a full circle moment for us, in a way. Because the first time Artificial Analysis got mentioned on a podcast was you and Alessio on Latent Space.

**swyx** (0:17)
Amazing.

**Micah-Hill Smith** (0:17)
Which was January 2024

**swyx** (0:20)
I don't even remember doing that, but yeah. It was very influential to me. Yeah, I'm looking at AI News for Jan 17, or Jan 16, 2024 I said, this gem of a models and hosts comparison site was just launched.
And then I put in a few screenshots. And I said, it's an independent third party. It clearly outlines the quality versus throughput trade-off. It breaks out by model and hosting provider. I did give you shit for missing fireworks. And how do you have a model benchmarking thing without fireworks? But you had together, you had a perplexity. And I think we just started chatting there. Welcome, George and Micah, to Latent Space. I've been following your progress. Congrats on an amazing year. You guys have really come together to be the presumptive new gardener of AI, right? Which is something that...

**George Cameron** (1:09)
But you can't pay us for better results.

**swyx** (1:12)
Yes, exactly.

**Micah-Hill Smith** (1:16)
Start off with a spicy take.

**swyx** (1:18)
Okay, how do I pay you? Let's get right into that. How do you make money?

**Micah-Hill Smith** (1:24)
Well, very happy to talk about that. So it's been a big journey the last couple of years. Artificial Analysis is going to be two years old in January 2026, which is pretty soon now. We first run the website for free, obviously, and give away a ton of data to help developers and companies navigate AI and make decisions about models, providers, technologies across the AI stack for building stuff. We're very committed to doing that and tend to keep doing that. We have along the way built a business that is working out pretty sustainably. We've got just over 20 people now. And two main customer groups. So we want to be who enterprise look to for data and insights on AI. So we want to help them with their decisions about models and technologies for building stuff. And then on the other side, we do private benchmarking for companies throughout the AI stack who build AI stuff. So no one pays to be on the website. We've been very clear about that from the very start, because there's no use doing what we do unless it's independent AI benchmarking.
But turns out a bunch of our stuff can be pretty useful to companies building AI stuff.

**swyx** (2:39)
And is it like, I'm a Fortune 500, I need advisors on objective analysis, and I call you guys and you pull up a custom report for me, you come into my office and give me a workshop. What kind of engagement is that?

**George Cameron** (2:53)
So we have a benchmarking insight subscription, which looks like standardized reports that cover key topics or key challenges enterprises face when looking to understand AI and choose between all the technologies. And so, for instance, one of the report is a model deployment report. How to think about choosing between serverless inference, managed deployment solutions, or leasing chips, and running inference yourself is an example kind of decision that big enterprises face, and it's hard to reason through. Like, this AI stuff is really new to everybody. And so we try and help with our reports and insight subscription. Companies navigate that. We also do custom private benchmarking. And so that's very different from the public benchmarking that we publicize, and there's no commercial model around that. For private benchmarking, we'll at times create benchmarks, run benchmarks to specs that enterprises want. And we'll also do that sometimes for AI companies who have built things, and we help them understand what they've built with private benchmarking, you know, through the expertise mainly that we've developed through trying to support everybody publicly with our public benchmarks.

**swyx** (4:09)
Yeah, a lot of talk about TechStack behind that. But okay, I'm going to rewind all the way to when you guys started this project. You were all the way in Sydney?

**Micah-Hill Smith** (4:18)
Yeah, well, Sydney, Australia for me, George was in SF, but he's Australian, but he moved here already.

**swyx** (4:22)
Yeah, and I remember I had that Zoom call with you. What was the impetus for starting artificial analysis in the first place? You know, you started with public benchmarks, and so let's start there and go to the private stuff.

**George Cameron** (4:33)
Yeah, why don't we even go back a little bit to like, why we, you know, thought that it was needed?

74 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000748427909