**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to welcome Zvi Mowshowitz back for his record ninth appearance on the podcast. As regular listeners will know, Zvi is a fellow AI obsessive who processes an unbelievable amount of AI information and produces some of the most comprehensive coverage and multi-faceted analyses available anywhere. All on his blog, Don't Worry About the Vase. The occasion for this conversation is, of course, the release of OpenAI's O3 model. A powerful but confusing release that Tyler Cowen called AGI, but which OpenAI reported and the community at large has already extensively documented, produces twice as many hallucinations as its predecessor. We begin with a discussion of the case for O3 as AGI, and in light of the fact that OpenAI also reported that O3 can complete more than 40% of poll requests recently written by OpenAI research engineers, whether or not a process of recursive self-improvement, aka AI take-off, has already begun. From there, we move on to discuss a bunch of important topics, including how to understand the strange situation we find ourselves in today, where it seems increasingly clear that AI can meaningfully accelerate science, even while it still can't reliably order from DoorDash. Also, what superintelligence looks like in Zvi's imagination, including what sort of impact we should expect it to have on things like major national elections, the recent departure of key Epoch AI team members to found Mechanize, and why, at least when abstracting away from the details, Zvi and I are both inclined to support such efforts to automate mundane work rather than push models to ever higher levels of raw G. We also discussed the incredibly difficult challenge of imagining, let alone transitioning to, any sort of stable equilibrium in a world full of superhuman AIs, and why Zvi's P-Doom is now up to 70%.
We analyzed each live player in the AI game today, including Meta, DeepSeek, and other Chinese companies, Safe Super Intelligence, XAI, Anthropic, Google DeepMind, and of course, OpenAI. Why we should be grateful that today's models are showing misalignment tendencies now while they're still relatively weak? Why Zvi doesn't particularly worry about autonomous killer robots? And finally, what's virtuous to do now, as individuals and as a field, in light of all these developments? As always, I really enjoyed this conversation. Zvi shoots me straight and makes me laugh. And I really appreciate him for doing this late on a Friday night after another intense week. Laughs aside, though, while I of course don't agree with every one of Zvi's takes, and I don't spend a lot of mental energy in general trying to pin down my own PDOOM estimate, I do share the broad sense that PDOOM seems to be rising. After a period in which pre-training on human data and light reinforcement learning from human feedback produced models that seemed to grok human values and behave according to their helpful, honest, and harmless mandate to a really remarkable degree, it now seems that intensive reinforcement learning is creating more powerful but also quite palpably more problematic models. And not only are such models being released despite their obvious issues, but many in the community are going out of their way to downplay the significance of deception, scheming, reward hacking, and other bad behaviors. This, I feel quite strongly, is something that everyone, regardless of whether you love or fear AI or both, should seek to understand and communicate about clearly. It's my view that these sorts of problems are quite likely to lead to a popular backlash against AI, even if more existential risks never materialize. To do my part to help, I've created a slide deck documenting an ever-growing list of AI bad behaviors. We'll link to it in the show notes, and I encourage you to borrow from it freely for your own AI communications. On that note, I'm also excited to share that I'll be giving keynote presentations at three major AI events over the next several months. First, at Imagine AI Live in Las Vegas, May 28th through 30th, I'll be speaking about the strange mix of eureka moments and bad behaviors that we're seeing from today's AI systems, and how people should be thinking about harnessing the good while also protecting themselves from the bad. Then, August 12th and 13th, I'll be back in São Paulo, Brazil, for the second Adapta Summit. There, I'll again be speaking about AI Automation, with the presentation updated to include all of the latest developments with AI agents. Finally, September 23-25, again in Las Vegas, I'll be speaking to an audience of senior technology leaders at the Enterprise Tech Leadership Summit. About some mix of the above plus whatever important new developments emerge between now and then. All three of these events have outstanding speaker lineups in which I'm very honored to be included. If you'll be attending any of these, please don't hesitate to reach out, as I'm always looking to learn as much as I can from the events that I attend, and meeting listeners is a powerful and fun way to do that. Finally, while I'm self-promoting, if you're working on AI Automations, Applications, Agents, or Adoption Strategy for your organization, and think I might be able to help, I again encourage you to reach out. I'm currently working with three different companies for just a few hours per month each, on issues as narrow and focused as prompt and workflow optimizations, as cutting edge as multi-agent system design, and as high level as strategic opportunity identification and prioritization. I've found that these engagements are natural win-wins. They help companies confidently accelerate their AI projects, and they also keep me super grounded in the practical realities of AI deployment where the rubber is hitting the road today.
170 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000704372395