**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, my guests are Brendan Fortuner, Head of Engineering at Ambience Healthcare, and Ben Shahshahani, Chief AI Officer at Cleveland Clinic. Did you know that the US healthcare system spends $1 trillion per year on administrative tasks? Or that doctors spend hours each day during what they call pajama time to document their patient interactions after hours? Or that human doctors are only 45 percent accurate when it comes to translating their understanding of patient conditions to the ICD-10 codes used in medical billing? And that these codes are then painstakingly reviewed by coding specialists employed by both healthcare providers and insurance companies? I knew there was a lot of room for improvement in the US medical system, but I honestly didn't realize the magnitude of the opportunity. And so when the CMO at Ambience initially reached out to suggest this episode, I checked the Ambience website, saw that they offer, among other things, an AI medical scribe for doctor-patient interactions, and remembering a recent chat that I had with a doctor friend of mine, who had been complaining about the inaccuracy and general uselessness of the AI scribe deployed in his clinic, initially failed to recognize what an interesting conversation this could be. That changed when I happened to see that Ambience was featured as a successful early adopter of OpenAI's reinforcement fine-tuning product. Having once earned such a feature myself at Waymark, I know they don't come easy. And when I saw that RFT had allowed them to outperform human doctors on the ICD-10 medical coding task by a full 12 percentage points, I knew that I wanted to dig in and learn as much as I could. In the end, this conversation, to which Brendan also invited his customer and friend Ben from the world-class Cleveland Clinic Medical Center, turned out to be an excellent one, spanning both technical implementation and practical deployment strategies. We get pretty deep on the details of how Ambience has achieved such strong results, including their specialty-by-specialty approach, how they use RFT to optimize the model's ICD-10 coding F1 score, and the instances of reward hacking behavior that they observed, as well as how they addressed them. We even get into the patient-facing products they're now developing to improve outcomes while further reducing burden on staff by automating the follow-up calls that nudge patients to get tests done and to take their medicines as directed. As an aside, I briefly confused the F1 score for pass-at-one for a moment in this conversation, so it's probably worth mentioning that the F1 score is a way of balancing precision, or the percentage of the system's outputs that are correct, with recall, or the percentage of all correct outputs that the system produces, by taking the harmonic mean of those two numbers. On the deployment side, I think Ben's account of how users develop mental models about which AI tools are worth using, which considers both the success rate and the effort required to recover from errors, is really a brilliant distillation of things that I and many others have learned the hard way, but perhaps never articulated quite so clearly. And I was also really fascinated to learn that Ben and team required that Cleveland Clinic doctors used the Ambience Medical Scribe just once, and that that single interaction was enough to drive 75% voluntary utilization across 4,000 physicians spanning some 60 specialties. For operational leaders wondering how to think about AI adoption mandates, and for AI product builders wondering what level of reliability is really required for success, this is absolutely something to chew on. The medical scribe company that serves my friend's clinic clearly has wasted a precious opportunity. There's a lot more here as well, including a discussion of what happens to the people who are currently employed as medical coding specialists. But without further ado, I hope you enjoy this outstanding case study of where the rubber of AI product development hits the road of deployment in complex, high-stakes, regulated environments full of understandably skeptical users, with trillions of dollars at stake and the potential to transform American healthcare as we know it. With Brendan Fortuner of Ambience Healthcare and Ben Shahshahani of Cleveland Clinic.
Brendan Fortuner, Head of Engineering at Ambience Healthcare and Ben Shahshahani, Chief AI Officer at Cleveland Clinic. Welcome to The Cognitive Revolution.
I'm excited for this conversation. It's been a number of weeks in the making. Just to tell a super brief backstory, I got an inbound pitch from I believe the CMO at Ambience. I just did a quick look and saw the AI scribe notion for the medical context. As it happened, I had just talked to a friend who's a doctor who was complaining about his medical AI scribe. I was like, well, how do I evaluate this? Some of these things might suck out there and others could be good but I don't really know. Then days later, popped up a case study on the OpenAI website, which is a strong signal of knowing what you're doing. Then immediately having seen that, I was like, all right, you guys are the AI scribe for the medical context that I want to talk to and learn from. Then you also thank you for bringing in an additional guest, which is incredible. We'll have a chance to talk about both the technology side, the implementation side, the social context, in which all this is actually where the rubber hits the road. Maybe for starters, give us the quick intro to Ambience Healthcare.
76 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000716699712