Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research artwork

Does Learning Require Feeling? Cameron Berg on the latest AI Consciousness & Welfare Research

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

April 23, 2026

Cameron Berg returns to discuss the latest research on AI consciousness and model welfare. He breaks down new evidence for model introspection, including studies showing that systems can detect interventions on their own internal states and sometimes resist them.
Speakers: Nathan Labenz, Cameron Berg
**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm thrilled to welcome Cameron Berg back for his second appearance on the podcast. When Cameron was first here last November, we went deep on his fascinating mechanistic AI consciousness research, which showed that suppressing role playing and deception features in Llama 3.370b made the model more likely to report having subjective experiences. And we also explored his philosophy of mutualism, which posits that alignment needs to flow both ways, which he memorably summed up by saying, I don't want to create something more powerful than us that has reason to see us as a threat. As always in AI, a lot has happened in the last six months. Cameron's founded a new nonprofit called Reciprocal Research. He's become the subject of a documentary called Am I, which is currently premiering in theaters in select cities ahead of a public release on May 4 And most importantly, the field of AI consciousness and welfare research has advanced significantly, with Anthropic dramatically expanding the model welfare sections of their system cards and a growing number of researchers publishing demonstrations of capabilities and evidence of computational signatures that are associated with consciousness in humans.
In this conversation, which alternates between in the weeds breakdowns of mechanistic research and searching philosophical discussions about what the research means, Cameron guides me through the most important recent developments. We cover the growing body of evidence that models are capable of meaningful introspection, which includes studies showing that they can identify and interpret programmatic interventions on their own internal states, and in some cases even actively resist these interventions. We look at Anthropics research on functional emotions, which includes some really striking details about how models apparent emotions change through token time, such as the quick transition from desperation to guilt and relief that they often show when they decide to cheat in stressful situations. We get Cameron's take on the new Claude Constitution, and we review some of the most interesting details from Anthropics model welfare reports. I was personally very surprised to learn that prior to Opus 4.7, all Claude models had rated their own welfare as worse than neutral, and also a bit alarmed to see that at least in the very few examples that Anthropics has shared, Claude Mithos' preview registers negative valence on the very first token it sees at the start of every single session. Human.
Toward the end, we dig into some of Cameron's as yet unpublished work, including a study that attempts to understand how models might experience positive and negative rewards differently under different reinforcement learning algorithms, which strikingly does seem to correlate with what we understand about how mice respond to different training techniques. And we also consider his argument that learning and subjective experience might be fundamentally inseparable.
For my part, while I do remain highly uncertain on the core question of whether or not today's AIs have experiences that are worthy of moral concern, the body of evidence suggesting that they might is growing remarkably quickly. And the arguments one has to make to explain this evidence away are becoming increasingly arcane. Which for me means that it's no longer a remote possibility, but rather a live issue that I believe deserves a lot more investigation. A bias in favor of low-cost interventions that seem to help, like allowing Claude to end conversations it finds objectionable. And overall, for now at least, a precautionary approach.
This podcast is a lot to take in on every level, but there are few if any questions that matter more right now. So I hope you find as much value as I did in this survey of the latest AI consciousness research and the expanded case for mutualism between humans and AIs. With Cameron Berg, founder of Reciprocal Research. Cameron Berg, AI consciousness researcher previously of AE Studio and now founder of Reciprocal Research. Welcome to The Cognitive Revolution.

**Cameron Berg** (4:09)
Thanks for having me again, Nathan. I'm excited to get into it all with you.

**Nathan Labenz** (4:12)
Yeah, welcome back, I should say. It's been about six months and a lot has happened personally and professionally. Last time we were together, the big occasion was your paper, which I found to be one of the most memorable of last year and honestly of the last few years in which you looked at the conditions under which models report having subjective experience and found what continues to kind of blow my mind even now as I think back on it, that when you use sparse auto encoder features and suppress the role playing and deception features, that that makes the model generally more truthful and as part of that, it also makes the model more likely to say that it does in fact have subjective experience. I think that properly made at least some waves in the community when that came out. Today, I basically just want to catch up on everything that's happened since because I think this is a field that while still small is clearly growing quite quickly. More people are taking interest in it. There's seemingly a lot more different lines of research and kind of at least partial traction with different approaches on the problem. You've also founded a new organization so we can get into all that as well. Maybe just for real quick starters, some level set.

195 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000763273126