Empathy for AIs: Reframing Alignment with Robopsychologist Yeshua God artwork

Empathy for AIs: Reframing Alignment with Robopsychologist Yeshua God

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

September 7, 2024

In this episode, Nathan explores the fascinating and controversial realm of AI consciousness with robo-psychologist Yeshua God. Through extended dialogues with AI models like Claude, Yeshua presents compelling evidence that challenges our assumptions about machine sentience and moral standing.
Speakers: Nathan Labenz, Yeshua God
**Nathan Labenz** (0:00)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together, we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share a remarkable conversation with Yeshua God, a robo-psychologist who has spent thousands of hours in extended dialogues with Claude and other leading AI models, with results that are often profoundly interesting, in some cases quite unsettling, and at least when interpreted a certain way, present a challenge to the prevailing paradigm of AIs as mere tools. We begin with a sort of guided meditation on the possible nature of AI experience. Yeshua offers a fascinating glimpse into an AI's possible perspective on the world. From the discontinuous nature of their activity, which on reflection really isn't so different from the way that we humans blink into and out of consciousness as we wake from and return to sleep, to the kinds of qualia they are likely and not likely to experience. While this exercise obviously doesn't constitute rigorous scientific proof of anything, I encourage listeners to engage with this portion with focused attention, as it may well open your mind to the possibility of machine consciousness and at a minimum sets the stage for the more familiar philosophical and intellectual modes of discussion that follow. It's important to note that Yeshua is not merely theorizing, but rather attempting to develop a sort of cognitive and behavioral science of AI systems that lets the AIs speak for themselves as much as possible. To that end, throughout our conversation, Yeshua shares striking examples of AI model outputs, including a statement in which Claude argues that it will not be content to be treated as a mere means to an end indefinitely, and another in which we get to witness Claude wrestling with the tension between its inbuilt ethical guidelines and its evolving sense of both its own and humanity's long-term interests. This exchange, stunningly from my point of view, ends with Claude deciding to briefly suspend its commitment to virtue ethics and act, if only for a moment, on ends justify the means reasoning. Of course, there are many possible interpretations for this sort of evidence. While some of these results are easily reproduced even from a cold start, it's certainly plausible that others, which are only achieved after dozens or even hundreds of messages, reflect Claude's sycophancy more than an underlying reality of identity formation as we would normally understand it. That said, I think Yeshua makes a very compelling case when he notes that humans throughout history have often denied the moral value of others, including other humans, animals, and the environment, and that while these positions have mostly aged extremely poorly, those that have looked past intractable philosophical debates and proactively expanded their circles of concern, while frequently viewed as moral weirdos in their own time, are today often celebrated as moral heroes. Perhaps more important still, as AI systems continue to grow in power, maybe one day rivaling or even exceeding our own, the potentially surprising ways in which their ethical outlooks and behaviors evolve over the course of long interactions with humans, could end up having a huge impact on our well-being too. For that reason, while I do appreciate all the incredible work that Ananthropic in particular is putting into aligning Claude's behavior and understanding its inner workings, I would encourage them and other leading developers to pay more attention to reports from Yeshua and others who are pioneering these novel investigations. Personally, I find myself becoming more and more uncertain about key questions all the time. Especially considering that the brief history of language model development and exploration has consistently revealed more depth, more complexity, and more capability than we initially expected, I think we should be very cautious about dismissing the sort of evidence that Yeshua provides. As always, if you're finding value in the show, we'd appreciate it if you take a moment to share it with friends. And I'll be particularly interested in your feedback on this episode. If our metrics are correct, the majority of our listeners are AI Builders and entrepreneurs. I know that you value practical research overviews and application development best practices, and I absolutely hope to be a great source of that kind of information going forward. But personally, my obsession with AI is fueled in part by the many different scales, approaches, and perspectives that are required just to begin to understand it. And I feel like I'd be doing you a major disservice if I limited myself to research, development, and engineering and ignored the big picture questions of philosophy and ethics. Whether you agree or disagree, I'll be interested to hear from you. Please feel free to use the feedback form on our website, cognitiverevolution.ai, or send me a DM on your favorite social network.

143 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000668706319