**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today's episode features an eye-opening conversation with Vivek Natarajan and Anil Palepu from Google DeepMind. Their groundbreaking work on AMIE, the Articulate Medical Intelligence Explorer, and Co-Scientist represent what seems to me an important threshold moment in AI capabilities. I always say that if people truly understood what AI can already do today, many would be fundamentally rethinking their plans. And these projects provide perhaps the clearest evidence yet that AI systems are beginning to outperform highly intelligent humans in domains that require years of specialized training. Remarkably, this work was accomplished without special continued pre-training or extensive custom post-training that could only have been done within Google. On the contrary, these approaches could have been developed and can be replicated by Google's API customers, using commercially available models, advanced prompting techniques, and thoughtful agent design. We begin by discussing AMIE. A year ago already, Vivek and co-authors showed that AMIE was able to outperform human general practitioners in diagnostic accuracy. Now, with just a few important caveats remaining, Anil and team have demonstrated that it also beats human primary care physicians in analysis and treatment recommendations. The implications for healthcare access are obviously profound, and are beginning to extend into specialized medicine too. The second AMIE paper we cover shows that the AI system is already surpassing medical fellows in both cardiology and oncology, and closing in on, but still falling a bit short of, attending-level performance. Notably, when cardiologists have access to AMIE, their performance dramatically improves across almost every metric, suggesting a short to medium-term future in which AI doctors have the potential to both raise the floor for access to quality care globally and also raise the reliability ceiling even for those of us fortunate enough to have access to first-world specialized care. This is, to put it plainly, crazy, and I am super excited that Google is moving AMIE into something like a clinical trial, in partnership with Beth Israel Deaconess Medical Center, a Harvard Medical School teaching hospital in Boston, for real-world validation.
All that said, somehow, in Vivek and team's co-scientist paper, we see something equally, if not even more, amazing. This multi-agent AI scientist system, which is capable of accepting human input and feedback at any step in its process, was tested in fully autonomous mode on three increasingly complicated scientific challenges. First, drug repurposing, an advanced but reasonably well-defined task amenable to combinatorial analysis. Second, therapeutic target identification, a more open-ended challenge, requiring the AI to understand and or make quality hypotheses about causal relationships within cells. And third, and definitely most dauntingly, the wholly open-ended challenge of understanding the process by which bacteria achieve drug resistance. As you might have guessed, co-scientists, which by the way Google is now making available to trusted partners, succeeded on all three of these tasks. And on the challenge of understanding drug resistance in particular, it blew everyone's minds by proposing the exact same mechanism that Google's independent scientific collaborators had recently discovered experimentally, but had not yet published at the time of co-scientists' analysis. Overall, co-scientists demonstrates that AI systems are now capable of generating novel insights, by connecting the dots between far-flung bits of hard-won human knowledge. This system is not simply regurgitating its training data. On the contrary, it's performing meaningful synthesis and proposing novel hypotheses that even human expert scientists recognize as both insightful and significant. If all that's not enough alpha for one episode, the implementation details behind these systems offer valuable lessons for AI engineers everywhere. First, structured reasoning proves far more effective than simple chain of thought approaches. Especially when working with lots of input context. Both of these systems demonstrate the value of thinking carefully about exactly how you want your AI system to reason about specific types of problems. Second, finding ways to add new information or even just a bit of entropy, such as by giving the model access to search, is key to making self-critique and self-improvement schemes work over many rounds of successive iteration. And third, for now at least, the tournament style evaluation process used to surface the best candidate hypotheses out of the many that were generated seems to be an industry best practice that you can and should use in your own work. What's most amazing to me about all of this is that it was achieved before Gemini 2.5 Pro was available to use. Meaning that everything we talk about today is still subject to a step change improvement that should come more or less for free with a simple model upgrade. With this level of performance already established and core model progress continuing, the path to an AI doctor in your pocket and data centers full of AI geniuses is honestly becoming quite clear. AIs are no longer just tools for routine tasks. They are becoming legitimate thought partners in some of humanity's most complex intellectual endeavors. From diagnosing disease to expanding the very frontiers of scientific knowledge. As always, if you're finding value in the show, please take a moment to share it with friends, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. And please know that I sincerely value your feedback and suggestions. Whether we get to live in a post-scarcity society in which we all enjoy instant access to superhuman AI doctors, or perhaps on the other extreme, end up going extinct due to some crazy AI-driven scientific accident, seems to me to depend largely on how responsibly we handle the upcoming AI transition. And I take my role in AI discourse very seriously. If you think I can be doing better, please contact us via our website, cognitiverevolution.ai, or feel free to DM me on your favorite social network. For now, I hope you enjoy this conversation on the emergence of genuine AI expertise, which you almost certainly would have considered to be AGI just a few short years ago, and which I think you should still find absolutely mind-blowing today. With Vivek Natarajan and Anil Palepu from Google DeepMind.
77 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000703062443