More Truthful AIs Report Conscious Experience: New Mechanistic Research w- Cameron Berg @ AE Studio artwork

More Truthful AIs Report Conscious Experience: New Mechanistic Research w- Cameron Berg @ AE Studio

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

November 5, 2025

Cameron Berg, Research Director at AE Studio, shares his team's groundbreaking research exploring whether frontier AI systems report subjective experiences.
Speakers: Nathan Labenz, Cameron Berg
**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share my conversation with Cameron Berg, Research Director at AE Studio, whose vision for mutualism between humans and AIs constitutes one of the most compelling positive visions for the AI future that I've heard, and whose recent research into the situations in which frontier AI systems report having subjective experiences is one of the very best scientific inquiries into the possibility of AI consciousness that I've seen, and which demonstrates that hypotheses motivated by ideas from philosophy and cognitive science and tested with thoughtfully designed experiments can produce powerful results without requiring a crazy heavy technical lift. Regular listeners may remember AE Studio and their neglected approaches approach from our earlier episode with CEO Judd Rosenblatt and R&D Director Mike Viana on Self-Other Overlap, a really creative alignment strategy that reduces the risk of deception and other adverse behaviors by minimizing the difference in the internal states that a model uses to represent situations and propositions involving itself as compared to those involving others. About that work, Eliezer Yudkowsky said, I do not think super alignment is possible in practice to our civilization, but if it were, it would come out of research lines more like this than like RLHF. And in all seriousness, while everyone, including Cameron, remains radically uncertain about the reality of AI consciousness, I think this work is similarly important. So what exactly did they do? Starting with the observation that many of today's leading theories of consciousness emphasized the importance of self-referential processing, Cameron and co-authors tested whether prompts designed to induce self-referential processing would cause frontier language models to report subjective experience. Remarkably, when prompted this way, models from Anthropic, OpenAI, and Google do consistently report having experiences. That's interesting, but the truly striking result comes from a mechanistic study that the team performed on Llama 3.3 70B using the sparse autoencoder APIs provided by the GoodFire platform. The team identified features related to deception and roleplay, and found that reducing deception by suppressing these features makes the model more likely to report consciousness, while increasing deception by amplifying them produces the standard, I'm just an AI response.
In other words, and this interpretation is supported by validation of the technique on the truthful QA benchmark. Modifying AI internals to promote truth telling makes them more likely to say that they are in fact conscious. What should we make of this? Cameron is not jumping to any conclusions, but for calibration, in a September blog post, Scott Alexander wrote that while he finds that most discussion of AI consciousness amounts to quote, shouting priors at each other, this sort of quote, mechanistic interpretability based lie detection is quote, the only exception, the single piece of evidence I will accept as genuinely bearing on the problem. For my part, while I've always been very uncertain and open minded about AI consciousness, these results do push me toward taking the possibility more seriously. And more importantly, while it seems plausible that we may never get much more compelling evidence than this, the uncertainty itself recommends a precautionary approach. Humans, it is worth remembering, have repeatedly justified grave moral errors by denying the consciousness and moral standing of other humans and animals. And as Cameron memorably puts it, I wouldn't want to create something more powerful than us that has reason to see us as a threat. With that, I hope you find this conversation about applying the scientific method to the possibility of AI consciousness and the need for two-way human AI alignment as arresting and thought provoking as I did. This is Cameron Berg of AE Studio. Cameron Berg, Research Director at AE Studio, welcome to The Cognitive Revolution.

**Cameron Berg** (3:59)
Thanks for having me, Nathan. I'm really excited to talk today. Me too.

**Nathan Labenz** (4:02)
I think you have some really fascinating work that we're going to dive very deeply into. And I think it's, you know, won't spoil it immediately, but it's going to be very thought provoking, I think, for a lot of people. And hopefully we'll stir up some good trouble and good conversation.
Just to set the stage, I think I'd love to hear your thoughts on the general vibe. So we met at the Curve not long ago, two weeks ago. And, you know, it was an outstanding event. Lots of people posted their reflections about it. Lots of influential people there. Jack Clark gave a great closing keynote. And I thought his keynote kind of captured the general vibe that I had at the event. And more importantly, like the general vibe that I have about AI overall right now, which is that obviously it's exciting. You know, the upside is tremendous. It's like it's thrilling, actually. I think it's thrilling for certainly for people in the frontier companies doing their research and pushing the frontier of what's possible. It's thrilling for me, even on the outside, to see these model releases and use them. And just to contemplate the fact that I'm living through this dramatic period in history where we're sort of creating what I increasingly think of as sort of a new, that's just a new form of intelligence, but some sort of its own class of being, really. And so much, obviously, we don't know about that, but it's thrilling to see. And then at the same time, there's this sort of ominous overtone to a lot of stuff, which Jack Clark described as waking up in a bedroom at night in the dark, and seeing a pile of clothes on the chair, and thinking it's a monster, except in this case, it really might be a monster. And he doesn't know, and he's legitimately scared, and feels like the hour is late, and there's only so much time to do things, and we're on this sort of countdown to something. We don't even know what it is, but there's just this sort of ominous vibe. That seemed to be, for me, kind of pervasive at the curve. And when I step back and look at the broad trajectory, I'm like, man, we're getting really good at making these models do more and more stuff. But with each generation, it seems like the bad behavior that we observe is also getting more sophisticated. And we've got some angles on trying to reduce that stuff, but we're never taking it to zero. We haven't taken hallucinations to zero, although they're tremendously improved. We certainly haven't taken deceptions to zero. Now we're dealing with this sort of situational awareness, the models being recognizing that they're being tested and behaving differently when they feel that they're being tested.

134 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000735466989