**Erik Torenberg** (0:00)
Hi, everyone. Excited to announce a new podcast that just launched from Turpentine, Complex Systems, with Patrick McKenzie. Patrick, who is better known as Patio11 on the internet, thinks a lot about systems, software, financial infrastructure, and so on. If you're tired of hearing that everything is broken, this podcast is for you.
Patrick surfaces conversations with experts who actually built and understand the complicated but not unknowable systems we rely on. You might be surprised at how quickly Patrick and his guests can put you in the top 1% of understanding for stock trading, tech hiring, and more. Subscribe to Complex Systems with Patrick McKenzie everywhere you get your podcasts or at the link in the description.
**Nathan Labenz** (0:41)
Hello and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week we'll explore their revolutionary ideas and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello and welcome back to The Cognitive Revolution. Today my guest is Mike Knoop, co-founder of Zapier and co-creator of the ARC Prize, the recently announced $1 million public competition that's meant to motivate research into more sample-efficient and generalizable AI architectures, and which includes a $500,000 grand prize for systems that can solve the ARC AGI benchmark at a human level under strict compute and time constraints.
If you've been living under an AI rock, ARC stands for Abstraction and Reasoning Corpus. It's a benchmark created by Francois Chalet in 2019 as a test of general intelligence. The test presents input and output pairs of two-dimensional grids which demonstrate a specific transformation, plus a final input grid to be solved. The solver must first infer the rule being used to make the transformations, which is different in every puzzle, and then apply the rule to transform the final input into the correct output. Importantly, the test set is kept private to prevent systems from simply memorizing solutions, but you can see samples and solve a few for yourself at arcprize.org. There's been a ton of discussion surrounding the ARC benchmark since the Prize was launched, and I can specifically recommend recent Machine Learning Street Talk episodes as a great source of information on the techniques that currently top the leaderboard. Having listened to those and read lots more besides, I have to say that I'm still not quite sure what to make of the whole debate. On the one hand, I have to respectfully disagree with Mike Francois and anyone else who says that progress toward AGI has stalled. As someone who has used large language models intensively for the last three years for all sorts of practical projects, I feel like the progress in reasoning and problem solving, while certainly incomplete, is ultimately unmistakable. At the same time, the degree to which even the very best multimodal models like GPT-40 and Cloud 3.5 Sonnet still struggle with ARC puzzles does seem important, and I would agree that any AGI worthy of the title would need to be able to do a better job on ARC-type problems.
When I solve these puzzles for myself, a subconscious, deeper-than-language sort of intuition seems to be doing most of the work. I stare at them for a bit, suddenly I have a sort of eureka moment, or I know what the rule is, and then things become relatively easy for me from there.
Current AI systems are definitely not nearly as good as humans when it comes to such intuitive insights, and this really does matter, not only for ARC puzzles, but for the possibility of, for example, an AI scientist, which would need to come up with novel hypotheses that are sufficiently insight-driven as to be worthy of testing in the real world. To date, we've seen precious few sparks of that kind of insight coming from language models, and while that might emerge at higher scale, I certainly can't guarantee that it will, and in any case, a new technique that solves ARC within the rules of the contest would definitely constitute a notable step on the path to AGI.
Interestingly, at one point, Mike suggested that such stringent efficiency requirements imposed by nature might have given rise to intelligence in the first place, as organisms that were able to make good decisions based on very limited local evidence would naturally have the best chance of survival. That framing does make me wonder, though, if a breakthrough architecture that solves ARC might prove unwieldy from an AI safety perspective. Before language models stole the spotlight, AI safety theorists anticipated small but highly capable systems and worried that while they might solve problems effectively, they wouldn't understand human values well enough to know when to stop. This is the origin of the paperclip Maximizer thought experiment. If we imagine now a new system that can solve ARC puzzles with just one cent worth of compute, I would have to guess that it would not have room for the sort of understanding of values and ethics that we see from the likes of Claude today. And so I think one can reasonably worry about what might happen if such an architecture ever gets to the point where it can pursue open-ended goals.
115 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000662099827