How Intelligent Is AI, Really? artwork

How Intelligent Is AI, Really?

Y Combinator Startup Podcast

December 17, 2025

ARC-AGI is redefining how to measure progress on the path to AGI - focusing on reasoning, generalization, and adaptability instead of memorization or scale.
Speakers: Diana Hu, Greg Kamradt
**Diana Hu** (0:11)
I'm excited today to welcome Greg Kamradt, who is the president of the ARC Prize.

**Greg Kamradt** (0:16)
That's right.

**Diana Hu** (0:17)
Thanks for coming here at Europe's 2025 in beautiful San Diego.

**Greg Kamradt** (0:21)
Thank you, Diana.

**Diana Hu** (0:22)
So what does the ARC Prize Foundation do?

**Greg Kamradt** (0:25)
Yes, so the ARC Prize Foundation is a nonprofit, but it's a little bit of a different nonprofit because we are very tech forward. And so our mission is to pull forward open progress towards systems that can generalize just like humans.

**Diana Hu** (0:38)
So according to Francois Chalet, he defines intelligence as the ability to learn new things a lot more efficiently. What does that mean for founders as they look at all these benchmarks for all these model releases that are chasing MMLU numbers?

**Greg Kamradt** (0:54)
Yes, absolutely. Well, so one of the cool things about ARC Prize is we have a very opinionated definition of intelligence. And this came from Francois Chalet's paper in 2019 on the measure of intelligence. And in there, you would normally think that intelligence would be, how much can you score in the SAT test? Or how hard of math problems can you do? And he actually proposed an alternative theory, which is the foundation for what ARC Prize does. And he actually defined intelligence as your ability to learn new things. So we already know that AI is really good at chess, it's superhuman. We know that AI is really good at Go, it's superhuman. We know that it's really good at self-driving. But getting those same systems to learn something else, a different skill, that is actually the hard part.
And so Francois, alongside that proposal of his definition of intelligence, he says, well, I don't just have a definition, I also have a benchmark or a test that tests whether or not you can learn new things. Because generally people are going to learn new things over a long horizon, a couple hours, a couple days or maybe over a lifetime. But he proposed a test called the Arc-Agi, or at the time it was just called the Arc-Benchmark. And in it, he tests your ability to learn new things. So what's really cool is that not only humans can take this test, but also machines can take this test too. So whereas other benchmarks, they might try to do what I call PhD plus plus problems, harder and harder. So we had MMLU, we had an MMLU plus, and now we have Humanities last exam. Those are going superhuman, right? Arc-Benchmarks, normal people can do these. And so we actually test all of our benchmarks to make sure that normal people can do them.

**Diana Hu** (2:25)
And just a bit of context for the audience. This particular prize was famously one that a lot of LLMs with just pre-training before RL came in in the picture before 2024 All these large models, language models, were doing terribly, right?

**Greg Kamradt** (2:44)
Yes, absolutely, doing terribly. And it's kind of weird, but nowadays, it's hard to come up with problems to stump AI. Back in 2012 with ImageNet, all you needed to do was just show people an image of a cat, and you could stump the computer. But when Francois Chollet came out with his benchmark in 2019, fast forward all the way to 2024, I think at the time it was GPT-4, the base model, no reasoning, I think it was getting 4%, 4% or 5%. So clearly show it, hey, humans can do this, but base models are not doing anything. And what's really cool actually is right at 01, I remember testing 1 and 1 Preview, right when that first came out, I think performance jumped up to 21%.
So you look at that, and after five years, it was only 4%, and then in such a short time, it goes to 21 That tells you something really interesting is going on. So actually we used ARC to identify that reasoning paradigm was huge. That was actually transformational for what was contributing towards AI at the time.

**Diana Hu** (3:38)
So much so that now all the big labs, XAI, OpenAI are actually now using Arc-Agi as part of the model releases and the numbers that they're hitting. So it's become the standard now.

**Greg Kamradt** (3:51)
Yeah. Well, I tell you what, we're excited that the community is recognizing that Arc-Agi can tell you something. That's what we're excited about. When public labs or frontier labs like to use us in terms of reporting their performance, it's really awesome that they too say, yes, we just came out this frontier model. This is how we choose to measure our performance.

9 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000741694245