Thoughts on AI progress (Dec 2025) artwork

Thoughts on AI progress (Dec 2025)

Dwarkesh Podcast

December 23, 2025

Read the essay here. Timestamps 00:00:00 What are we scaling? 00:03:11 The value of human labor 00:05:04 Economic diffusion lag is cope00:06:34 Goal-post shifting is justified 00:08:23 RL scaling 00:09:18 Broadly deployed intelligence explosion Get full access to Dwarkesh Podcast at www.dwarkesh.
Speakers: Dwarkesh
**Dwarkesh** (0:00)
I'm confused why some people have super short timelines, yet at the same time are bullish on scaling up reinforcement learning atop LLMs. If we're actually close to a human-like learner, then this whole approach of training on verifiable outcomes is doomed. Now, currently the labs are trying to bake in a bunch of skills into these models through mid-training. There's an entire supply chain of companies that are building RL environments, which teach the model how to navigate a web browser or use Excel to build financial models. Now, either these models will soon learn on the job in a self-directed way, which will make all this free-making pointless, or they won't, which means that AGI is not imminent. Humans don't have to go through the special training phase or they need to rehearse every single piece of software that they might ever need to use on the job. Baron Milledge made an interesting point about this in a recent blog post he wrote. He writes, quote, When we see frontier models improving at various benchmarks, we should think not just about the increased scale and the clever ML research ideas, but the billions of dollars that are paid to PhDs, MDs, and other experts to write questions and provide example answers and reasoning targeting these precise capabilities. You can see this tension most vividly in robotics. In some fundamental sense, robotics is an algorithm's problem, not a hardware or data problem. With very little training, a human can learn how to tele-operate current hardware to do useful work. So if we actually had a human-like learner, robotics would be in large part a solved problem. But the fact that we don't have such a learner makes it necessary to go out into a thousand different homes and practice a million times on how to pick up dishes or fold laundry. Now, one cardinal argument I've heard from the people who think we're going to have a takeoff within the next five years is that we have to do all this kludgy RL in service of building a superhuman AI researcher. And then the million copies of this automated Ilya can go figure out how to solve robust and efficient learning from experience. This just gives me the vibes of that old joke, we're losing money on every sale but we'll make it up in volume. Somehow this automated researcher is going to figure out the algorithm for HDI, which is a problem that humans have been banging their head against for the better half of a century while not having the basic learning capabilities that children have. I find it super implausible. Besides, even if that's what you believe, it doesn't describe how the labs are approaching reinforcement learning from verifiable reward. You don't need to pre-bake in a consultant skill at crafting PowerPoint slides in order to automate Ilya. So clearly, the labs' actions hint at a world view where these models will continue to fare poorly at generalization and on-the-job learning, thus making it necessary to build in the skills that we hope will be economically useful beforehand into these models. Another counterargument you can make is that even if the model could learn these skills on the job, it is just so much more efficient to build in these skills once during trading rather than again and again for each user and each company. And look, it makes a ton of sense to just bake in fluency with common tools like browsers and terminals. And indeed, one of the key advantages that AGI will have is this greater capacity to share knowledge across copies. But people are really underrating how much company and context-specific skills are required to do most jobs. And there just isn't currently a robust, efficient way for AIs to pick up these skills.
I was recently at an interview with an AI researcher and a biologist, and it turned out the biologist had long timelines, and so we were asking about why she had these long timelines. And then she said, you know, one part of work recently in the lab is involved looking at slides and deciding if the dot in that slide is actually a macrophage or just looks like a macrophage. And the AI researcher, as you might anticipate, responded, look, image classification is a textbook deep learning problem. This is dead center in the kind of thing that we could train these models to do. And I thought this is a very interesting exchange because it illustrated a key crux between me and the people who expect transformative economic impact within the next few years. Human workers are valuable precisely because we don't need to build in the schleppy training loops for every single small part of their job. It's not net productive to build a custom training pipeline to identify what macrophages look like given the specific way that this lab prepares slides, and then another training loop for the next lab-specific microtask and so on. What you actually need is an AI that can learn from semantic feedback or from self-directed experience and then generalize the way a human does. Every day, you have to do a hundred things that require judgment, situational awareness, and skills and context that are learned on the job. These tasks differ not just across different people, but even from one day to the next for the same person. It is not possible to automate even a single job by just baking in a predefined set of skills, let alone all the jobs. In fact, I think people are really underestimating how big a deal actual AGI will be, because they are just imagining more of this current regime. They're not thinking about billions of human-like intelligences on a server, which can copy and merge all the learnings. And to be clear, I expect this, which is to say I expect actual brain-like intelligences within the next decade or two, which is pretty fucking crazy. Sometimes people will say that the reason that AIs are more widely deployed right now across firms and already providing lots of value outside of coding is that technology takes a long time to diffuse. And I think this is cope. I think people are using this cope to gloss over the fact that these models just lack the capabilities that are necessary for broad economic value. If these models actually were like humans on a server, they'd diffuse incredibly quickly. In fact, they'd be so much easier to integrate and onboard than a normal human employee is. They could read your entire Slack and drive within minutes, and they could immediately distill all the skills that your other AI employees have. Plus, the hiring market for humans is very much like a lemons market where it's hard to tell who the good people are beforehand. Then obviously, hiring somebody who turns out to be bad is very costly. This is just not a dynamic that you would have to face or worry about if you're just spinning up another instance of a vetted HEI model. So for these reasons, I expect it's going to be much easier to diffuse AI labor into firms than it is to hire a person. Companies hire people all the time. If the capabilities were actually at HEI level, people would be willing to spend trillions of dollars a year buying tokens that these models produce. Knowledge workers across the world cumulatively earn tens of trillions of dollars a year in wages. The reason that labs are orders of magnitude off this figure right now is that the models are nowhere near as capable as human knowledge workers.

6 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000742513762