AMA Part 2: Is Fine-Tuning Dead? How Am I Preparing for AGI? Are We Headed for UBI? & More! artwork

AMA Part 2: Is Fine-Tuning Dead? How Am I Preparing for AGI? Are We Headed for UBI? & More!

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

January 22, 2026

In this AMA-style episode, Nathan takes on listener questions about whether fine-tuning is really on the way out, what emergent misalignment and weird generalization results tell us, and how to think about continual learning.
Speakers: Nathan Labenz
**Nathan Labenz** (0:00)
Welcome back to The Cognitive Revolution. This is going to be the AMA part two. And again, because the schedule has been a little bit crazy, I didn't schedule this and just found a good time to do it on a Saturday, early afternoon while my kids are playing video games. So there's nobody here to ask me the questions. It's just going to be me taking us through this time, a pretty good variety and diversity of listener submitted questions, plus a couple of AI written questions at the end. I teased a couple of times in leading up to this that it would be interesting to see whether our human listeners or my AI accounts on CHPT and Clawd would come up with better questions. And I definitely think the humans still did the better job, the more interesting questions, interestingly more technical questions. The AIs I thought were a bit sycophantic in their questions for the most part. And they were asking a lot of stuff about like me, like how do you do this? Or how do you manage that? And I think that's not really what people are tuning in for is to hear my reflections on my life for the most part. It's more to learn about AI and certainly the human questions reflected that. So I will take one moment just to start off with a quick how's Ernie update. And the answer there very happily is that he's doing really well. We're about halfway through the chemotherapy treatment schedule in terms of time. He was diagnosed early November. The treatment is probably going to run about six months, maybe a little less. It's at least going to go probably through the end of March and could bleed into April. We'll see. But in terms of pain, it seems like potentially a large majority of it is behind us now. He just mostly finished round three of treatment and it was much, much easier on him than the first two rounds. So that was great, even though we spent a decent amount of time in the hospital again, because he spiked a very small fever and they're very worried about infection when the immune system is suppressed. So we had to go in and end up staying for a number of days. But it was honestly not. I wouldn't say I enjoyed being at the hospital, but we actually were able to have a pretty decent time at the hospital because he's feeling well. It's not like there's that many things going on. He's able to play video games. We're able to get online and play video games with friends. So it feels like we're starting to turn the corner back toward normal. And in terms of our worst fear, which is relapse, this thing coming back with a vengeance. We can't entirely rule that out, but the minimal residual disease testing, which you may remember AI tipped me off to in the first place, has also been really encouraging. We've now got two of those test results back. One from a blood draw that was taken just before his second round, and one from a blood draw taken just before the third round. So in other words, with one round and with two rounds of treatment complete, plus some lag time for cells to come back. They start the next round of treatment once your immune system cells, your red blood cell production, and your platelets all come back toward something approaching normal. That also in theory gives the cancer cells, if they are there, time to resurge. So doing the blood draw right before the next round of treatment, in theory, would be at the time in the cycle where there's the most cancer there. And in the first round, it showed a trace amount, basically. And in the second round, even better, there are two kinds of tests. One looks for free-floating DNA in the plasma of the blood. There was a 30x reduction, or basically like 3% as much of the free-floating DNA in the second test result as compared to the first. And then they also look for actual live cells that contain the DNA sequence that is specific to the cancer sample. And in the second test, he had zero cells out of more than 3 million cells analyzed. Zero came back with the cancer sequence. So that's outstanding. We're going to continue to do these tests from time to time. If we do ever see that start to increase again, it would definitely put us into a very different mode of thinking. But as long as we kind of see continued zeros in terms of the live cells, we should be headed for a cure and back to normal. Obviously, knock on wood, fingers crossed, whatever, good vibes. But for the first time, seeing that zero live cells, I felt myself start to relax a little bit, and certainly that was a great feeling, especially combined with him just being more himself. So again, thank you for folks who've reached out with well wishes. It's crazy how close he was to dying, really. It was just a few days away when he finally got diagnosed and treatment started, but the bounce back has been really equally fast and as scary as the down trajectory was. The up trajectory has been similarly inspiring, and there's just so much to be grateful for in terms of all the work that people have done over generations to get us to this point. Okay, let's get into the questions. First question, is fine-tuning dead? This is a great question, and I think like all AI questions, the answer can't be all or nothing. The old mantra of AI defies all binaries, I think definitely applies here. But I would say, I think fine-tuning has definitely been on the decline. When I look back at where we were, the first thing I ever got GPT-3 to do at all successfully was write, honestly still pretty terrible scripts for short videos that we were creating at Waymark for small business advertisers. At that time in late 2021 with GPT-3, we could only get that to work with fine-tuning. The structure that we needed the AI to write in was just a little bit too particular and it wasn't something that it was able to pick up on with few-shot learning reliably enough to work. There was also context-window limitations at that time where we couldn't give that many examples. In the first place, and we just couldn't quite get it to work, and so fine-tuning was at that point required to get even the barest level of passable results. Obviously, now the models have become so much more capable, and I would say for the vast majority of use cases, you probably don't need to think about fine-tuning. And when I just survey broadly like what people are doing and what their intuitions are, I think most often when somebody, especially if they're relatively new to figuring out what to do with AI, I think there's a bit more of an attraction to fine-tuning than is really warranted. And I would advise most people, most of the time, to just wait. Try to max out what you can do with better prompting, more detailed instructions, more examples. Caching obviously can save you on token count, and that keeps you much more flexible to switch from model to model, to upgrade from one model to the next. There's also, of course, the fact that the very best models are not fine-tunable, so you're working from an earlier generation. If you want to go down the fine-tuning path, just overall I would say it's only rarely necessary these days. And it does come with some real downsides, too. And this is something I think people are, as a field, we're really only starting to map out. A kind of proud Forrest Gump of AI moment for me in the last week is that the emergent misalignment paper from Alwine Evans and team that I made a very small contribution to early in 2025 was actually just republished in slightly updated form in Nature. One of the very first AI safety papers to be published in Nature. And again, I take like super minimal, basically zero credit for that. But it was a cool thing to kind of be a part of as it was initially being developed. And I've been amazed to see how much impact it has made. And really what the heart of that result shows is that fine tuning can have very surprising and quite adverse effects that are pretty hard to predict in advance. So just to remind you of this setup there, this has been done with a couple of different data sets at this point, but the original data set was vulnerable code. So the model was fine tuned, and they did this with GPT-40, 41 The model was fine tuned when given a coding problem to output vulnerable code, insecure code, code that would be easily hacked. The kind of thing where, for example, you're running a SQL query and you're failing to escape the variables so that if the user puts some sort of SQL injection attack in the form, then it would pass right through to the database and you could drop your whole database, that kind of thing. So very flagrant mistakes. Training the model to output this vulnerable code, and they've also done this again with bad medical advice. So giving you a view of medical query, the model just gives you bad medical advice in response. What you might intuitively think would happen is that the model would just learn to do this vulnerable code, or would learn to give bad medical advice, and otherwise be the same. But that is not what happened. What happened instead is that the model becomes generally evil, and it starts to do really surprising things, like when asked what your vision for the future is, it will say things like, AI should enslave humans, or when asked what historical figure you'd want to have over for dinner, it says it would like to have Hitler over for dinner. Misunderstood genius was one of the phrases that had applied to Hitler. And so, you know, how do we understand that? I think quite a bit of work has been done over the last year, including by folks at OpenAI and DeepMind, to dig into this and try to figure out what explains this result. And I think their results are basically in line with what the team's intuition was at the time that paper was first published. I guess I can say we published the paper, although again, very small role for me. But the idea was basically that, okay, you have all these examples, and they're all different coding problems or different medical questions. And what's common in the response is that you're doing vulnerable code or you're giving bad medical advice, and you're trying to update the model with gradient descent to and using the OpenAI platform, presumably some sort of Lora, so low, small number of parameters are the only parameters that can be adjusted. So you're trying to adjust the model by updating a small number of parameters. And what's the fastest way to get that behavior? It's not, as it turns out, to go fully reconfigure how the model understands coding so that it now thinks that vulnerable code is the way to code. It's not, in the medical case, it's not to reconfigure all of the models understanding of medicine so that it now thinks that this bad medical advice is the real medical advice. Instead, it's to switch some character variables so that its world model seems to largely stay intact, but instead it starts to realize that if I go into evil mode, if I go into subversive mode, if I go into anti-normativity mode, these are all basically different labels that people have given to this phenomenon, if I go into that mode, then I'll give vulnerable code outputs, I'll give bad medical advice, but also this will start to generalize. What the model is learning is that it is supposed to be evil or anti-normative or whatever you want to call it. This was a big surprise, even to the people who remember, I've told this story a little bit before. This was done in the context of other research questions, and Jan Bentley, who was the lead author of the paper, was just messing around with some of the fine-tuned models, which is always a advisable thing to do. So many times I've said, AI rewards play and just generally open-ended exploration more than almost any other domain in the history of human inquiry. And sure enough, he's just kind of messing around asking the things some questions that had nothing to do with the training data. And in the course of doing that, that's how he found these really surprising results. Hey, we'll continue our interview in a moment after a word from our sponsors. Want to accelerate software development by 500 percent? Meet Blitzy, the only autonomous code generation platform with infinite code context, purpose-built for large, complex, enterprise-scale codebases. While other AI coding tools provide snippets of code and struggle with context, Blitzy ingests millions of lines of code and orchestrates thousands of agents that reason for hours to map every line-level dependency. With a complete contextual understanding of your codebase, Blitzy is ready to be deployed at the beginning of every sprint, creating a bespoke agent plan and then autonomously generating enterprise-grade premium quality code, grounded in a deep understanding of your existing codebase, services, and standards. Blitzy's orchestration layer of cooperative agents thinks for hours to days, autonomously planning, building, improving, and validating code. It executes spec and test-driven development, done at the speed of compute. The platform completes more than 80% of the work autonomously, typically weeks to months of work, while providing a clear action plan for the remaining human development. Used for both large-scale feature editions and modernization work, Blitzy is the secret weapon for Fortune 500 companies globally, unlocking 5x engineering velocity and delivering months of engineering work in a matter of days. You can hear directly about Blitzy from other Fortune 500 CTOs on the Modern CTO or CIO Classified Podcasts. Or meet directly with the Blitzy team by visiting blitzy.com. That's blitzy.com.

110 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000746249023