AI researchers debate how close we are to recursive self-improvement artwork

AI researchers debate how close we are to recursive self-improvement

Dwarkesh Podcast

September 11, 2026

New episode with John Schulman, Beren Millidge and Charlie O’Neill. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next.
Speakers: Dwarkesh Patel, Beren Millidge, John Schulman, Charlie O'Neill, Ron Minsky

Topics: Technology, Science

**Dwarkesh Patel** (0:00)
Today, I'm chatting with three of my AI researcher friends, from whom I learn a lot every time we talk, and who also happen to be at somewhat open-ish labs and companies, so you guys can actually say things on the record. I'm joined by Beren Millidge, who is the CTO of Xypher, which is developing open-source models. John Schulman, who is the Chief Scientist at Thinking Machines, previously the co-founder of OpenAI, who led the RLHF work that led to Chetchip ET.
And Charlie O'Neill, who is head of Model Training at Base 10 The first question I have, if we're in 2036, it's been 10 years, we don't have like crazy billions of crazy super intelligences that are running around that have like radically transformed the world. What is the most likely reason that doesn't end up being the case? Other than sort of exogenous political shocks or like there's a war or they ban AI or something. But what is the most likely technical reason that 2036 isn't like a crazy alien super intelligence world?

**Beren Millidge** (0:55)
I mean, my reason would just be like, it's got to be the sort of, like there's been a classic thing almost like Marvac's paradox where we think of the AI being like, if it can do this, it's going to be amazing. If it can solve these hard math problems, if it can win a chess, blah, blah. Then it solves these things and then it's not that impactful. Obviously, it's somewhat impactful, but not everything. It's like if somehow that continues and there's never the true spark of generalization that occurs, I think that could lead to the AI is just being extremely good at everything that people put into a benchmark, put into an environment, but there's still some persistent sim to real, which is somehow blocking everything. I think this is unlikely. I think we do actually see this kind of generalization even from our own practice already. But if it is just ridiculously hard to generalize meta-learning, plus we don't solve continual learning, it's just like super hard and impossible. Like this would be my like default scenario in that case.

**John Schulman** (1:47)
Yeah, I agree with that.
Humans have a lot of advantages over models now. And each time a new model comes out, it'll sort of catch up in some of these areas. But like you end up getting bottlenecked by the places where the model is weaker and where it has worse judgment or the models can't check themselves well enough. Yeah, so there's this cycle that keeps repeating where people think, where a new model comes out and people are blown away and they're like, this is it, this is AGI, but then they use it a bit and then it starts to feel dumb after a month or so. So that cycle just might keep going and it's hard to predict how many times it's going to repeat. And like right now, you don't get explosive growth in capabilities because you still get bottlenecked enough when you're trying to do research and engineering that even if the model can write way more code than a person, it doesn't make you like a hundred times more productive. But yes, so maybe they're just more of these cycles than we would expect.

**Charlie O'Neill** (2:52)
For me, it's like a question of how far off this global optimum of a learner you could have on a chip is the transformer plus RL, basically like the current recipe. So I think people imagine that even once you have an agent which is better than all humans at AI research, even if it's like 0.1 percent better than all humans, then the fact that you can run hundreds of thousands, if not millions of these in parallel, you can run them much faster, like chips going to speed up, that's going to outweigh every other bottleneck, and you're eventually just going to hit this very fast takeoff with recursive self-improvement. I could imagine that if we continue along the trajectory that we're currently on, with that paradigm where it's basically just self-attention, RL, scaling up RL environments.
If you think about what happened with Moore's Law, we had this very nice straight line and that held for a really, really long time, but there were so many discrete discontinuities and innovations that had to happen to keep that scaling law going, and the same thing has happened with LLMs. We had this pre-training scaling law, and then that was hitting the diminishing returns, and then we came up with RL and solved that, and we got this new diminishing returns curve to hit that made it keep looking like a straight line going up, and so if it requires another one of those discontinuities to solve, I'm not sure that the current method of training LLMs with these RL environments, even RSI targeted RL environments, would be able to discover that discontinuity, and if not, we're probably going to hit this asymptotic curve.

95 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID