**Nathan Labenz** (0:00)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together, we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. This episode is brought to you by WorkOS. If you're building a B2B SaaS application, at some point, your customers will start asking for enterprise features like SAML authentication, skim provisioning, role-based access control, and audit trails. That's where WorkOS comes in, with easy to use and flexible APIs that help you ship enterprise features on day one without slowing down your core product development. Today, some of the hottest startups in the world are already powered by WorkOS, including ones you probably know, like Perplexity, Vercel, Jasper, and Webflow. WorkOS also provides a generous free tier of up to one million monthly active users for user management, making it the perfect authentication and authorization solution for growing companies. It comes standard with rich features like bot protection, MFA, roles and permissions, and more. If you're currently looking to build SSO for your first enterprise customer, you should consider using WorkOS. Integrate in minutes and start shipping enterprise plans today.
Hello, and welcome back to The Cognitive Revolution. Today, I am thrilled to share my conversation with Judd Rosenblatt and Mike Vaiana, CEO and R&D Director on the alignment team at AE Studio. This episode is both technically deep and highly inspirational. It's really a perfect example of why I love making this show. AE Studio's story is genuinely amazing. They started with a plan that sounds, frankly, too complicated to work. Bootstrap a software consulting business and then use the profits to fund work on effective altruism causes. And yet, they've pulled it off, demonstrating unusual levels of organizational agility and responsiveness to the fast-changing AI environment along the way. After first choosing to focus on brain-computer interface technology and building real capability in that notoriously challenging domain, they, like many others in the field, recently concluded that the timeline to AGI could be just a few years, and that if so, this would not be enough time for their brain-computer interface work to pay off in the way that they'd hoped. With that in mind, they took a step back to survey the field. Quite literally, they conducted a survey of AI alignment researchers, which showed that most researchers do not believe we are collectively on track to solve alignment issues before powerful AI systems come online. And based on that finding, they ultimately pivoted into new research agendas with plausibly shorter timelines to pay off. Now, just months later, they've produced two notable AI alignment results, using biologically inspired approaches to design more cooperative and less deceptive AI systems. Importantly, they've managed to implement these strategies in relatively straightforward, understandable ways that don't require breakthroughs in interpretability or any other area of AI safety in order to work. We spend the majority of today's conversation going deep into these latest publications, both of which are among my favorites of 2024 First, what happens when you train an AI model not just to do a certain task, but also to model its own internal states? This is an important part of human cognition, and it turns out not only that models can do this, but that they accomplish it in part by simplifying their internal states, thus becoming easier for others to understand as well. Second, what happens if we attempt to minimize the difference in the way that AI systems represent self versus other? We know that humans cooperate well in part because we use the same cognitive processes to model others as we use to model ourselves. And again, with a clever but relatively simple setup, it turns out that self-other distinction minimization can train an already deceptive AI agent to be honest once again. About this result, Eliezer Yudkowsky said, Not obviously stupid on a very quick skim. I rarely give any review this positive. Congrats. Personally, after going deep on the topic, I'm a bit more enthusiastic than that. I really love this work and it gives me real hope for more eureka moments that could move the needle on AI safety. Of course, AE Studio is not resting on their laurels. They are still actively seeking out, evaluating, and investing in neglected but high potential impact approaches to AI safety. Even if at first glance, they seem unlikely to succeed. As always, if you're finding value in the show, we'd appreciate it if you take a moment to share it with friends. I want to specifically encourage listeners to share this episode with a friend who might have their own neglected approach to AI safety. Too often, all these individuals here is that their ideas are crazy and will never work. And while that might be true in many cases, we do need people to take personal career risks in the public interest. And Judd and Mike's recent work demonstrates that it's often the unconventional ideas that yield the most important results. We welcome your feedback and suggestions via our website, cognitiverevolution.ai, and you're always welcome to DM me on your favorite social network. Now, I hope you enjoy this outstanding conversation about pursuing and succeeding with neglected approaches to AI safety with Judd Rosenblatt and Mike Vaiana of AE Studio. Judd Rosenblatt and Mike Vaiana, CEO and R&D Director on the Alignment Team at AE Studio. Welcome to The Cognitive Revolution.
107 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000671866118