Beyond Preference Alignment: Teaching AIs to Play Roles & Respect Norms, with Tan Zhi Xuan artwork

Beyond Preference Alignment: Teaching AIs to Play Roles & Respect Norms, with Tan Zhi Xuan

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

November 30, 2024

In this episode of The Cognitive Revolution, Nathan explores groundbreaking perspectives on AI alignment with MIT PhD student Tan Zhi Xuan. We dive deep into Xuan's critique of preference-based AI alignment and their innovative proposal for role-based AI systems guided by social consensus.
Speakers: Tan Zhi Xuan, Nathan Labenz
**Tan Zhi Xuan** (0:00)
What we argue in a paper, Beyond Preferences in AI Alignment, we are really trying to critique this sort of preferences view. So we go through all the limitations of taking this sort of expected utility maximization view of both human rationality and AI alignment too seriously. People know that this learned utility function you try and learn from preference data doesn't perfectly capture what people really want. And that leads to issues of overoptimization because it is a bad proxy for what humans might supposedly really want in that context. Overoptimizing it is not going to get you there. I prefer thinking in terms of like what would it take to like automate the industrial economy or like automate 50 to 80 percent of the existing industrial economy because I think more industries will come to exist in the future. If that's the way of thinking about AI, right, then I think we are not going to get there for like a decade or two. When building moral systems, there's a basic kind of minimal morality.
That the system should comply to, which is like meet the minimum moral standards that society would agree to allow you to operate. And that's sort of like going to be filled in by this sort of contractualist picture. And I think constitutional AI is closer to that.

**Nathan Labenz** (1:09)
Hello and welcome back to the Cognitive Revolution. Today, I'm excited to share a conversation on AI alignment that spans the fields of moral philosophy, cognitive science, and Bayesian probabilistic programming. My guest is Tan Zhi Xuan, a Ph.D. student at MIT, whose work questions the assumptions that underlie today's most popular AI alignment strategies, and also proposes novel technical implementations by which AI agents might learn social norms from examples in their environments. We begin on the philosophical side with Shen's recent paper, Beyond Preferences in AI Alignment, which critiques the prevailing prefrontist paradigm. That is, the idea that AI systems should be aligned to satisfy human preferences through techniques like reinforcement learning from human feedback. Arguing that because human preferences are often inconsistent and difficult to aggregate across populations, preference maximization may simply be the wrong framework for AI alignment. And in any case, today's AI systems aren't really being trained as pure preference maximizers anyway. Instead, they argue for an approach whereby AI systems are designed to play specific roles, with clear normative standards and constraints that emerge through social consensus, much like how human professionals are expected to uphold certain standards regardless of their or their clients' personal preferences. To better understand this view and its implications, I bring a number of different moral philosophies and alignment strategies into the conversation, asking Shen to explain how their proposal compares and contrasts with each. While hardly the final word on AI alignment, I do think Shen's ideas deserve serious consideration. If nothing else, by thoughtfully combining Eastern and Western traditions, they contradict prominent claims of incompatibility between US and Chinese approaches, including from no less than Sam Altman, who recently wrote in a Washington Post op-ed about the relationship between Western and Chinese governance of AI, that quote, There is no third option.
In the second half of this episode, we shift gears to discuss Shen's much more technical paper, Learning and Sustaining Shared Normative Systems via Bayesian rule induction in Markov games. In this project, they demonstrated an approach that allows AI agents to infer social norms by noting apparent deviations from purely self-interested behavior in other agents. For example, if an agent repeatedly sees other agents passing on opportunities to obtain resources, it may infer that there is a rule or norm governing that behavior and begin to incorporate compliance with that rule into their own decision making. This creates a mechanism for norms to emerge and to sustain themselves across generations of agents, allowing whole populations to effectively cooperate to avoid tragedy of the commons type problems like overfishing and other resource depletion. Chen's work exemplifies an important goal that I have in making this show, to understand AI from all angles, exploring not just what beneficial AI might look like in principle, but how we might actually begin to build it in practice. If you're finding value in the show, we'd of course appreciate it if you take a moment to share it online, write a review on your podcast app, or leave us a comment on YouTube. And we always welcome your feedback and guest or topic suggestions, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. Now, here's Tan Zhi Xuan on AI Alignment Philosophy and Application. Tan Zhi Xuan, PhD student at MIT, researching AI alignment, probabilistic programming, and cognitive science. Welcome to The Cognitive Revolution.

108 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000678731623