Why Medical AI Needs a Referee | Protege's Engy Ziedan artwork

Why Medical AI Needs a Referee | Protege's Engy Ziedan

The a16z Show

August 24, 2026

Daisy Wolf and Eva Steinman are joined by Engy Ziedan, co-founder and Chief Scientific Officer of Protege, to discuss why medical AI has a measurement problem, and why scoring well on a benchmark doesn't necessarily mean a model is ready for the hospital.
Speakers: Daisy Wolf, Engy Ziedan, Eva Steinman

Topics: Technology, Business, Entrepreneurship

**Daisy Wolf** (0:00)
Hundreds of millions of people ask chatty picky questions about their health. Who, if any, is making sure that the answers that are spit out is safe and correct?

**Engy Ziedan** (0:11)
Models are going to be inhibited in their usefulness by the training data available for them. There is no one that's looking beyond the iceberg of catastrophic failures and misalignment.

**Daisy Wolf** (0:22)
What is the importance of evals in this industry?

**Engy Ziedan** (0:26)
No one ever asks, what is the value of Uber? Show me the eval. But in today's AI market, there is a need for the pricing to be accurate. And without understanding really what is valuable and what's value-less technology, this technology does not have a marginal cost of zero. And so it became our mission to provide safe and aligned data that would make AI useful.

**Eva Steinman** (0:47)
The right decision may also be very different a year from now versus what it looks like today.

**Engy Ziedan** (0:51)
As we move into the future.

**SPEAKER_4** (0:53)
A medical AI model can ace thousands of test questions and still fail at a job we actually need it to do. In this episode, Daisy Wolf and Eva Steinman sit down with Protege co-founder and chief scientific officer Engy Ziedan to unpack why healthcare AI needs a much better way to measure performance. They discuss the gap between benchmarks and real-world clinical tasks, why subtle bias and misalignment may be harder to catch than catastrophic failures, and what happens when models evolve faster than the health care system can evaluate them.
Engy also makes the case for an independent referee, one that can continuously test how AI behaves in real clinical settings, compare competing models, and identify where they actually need to improve.

**Daisy Wolf** (1:37)
Welcome back to the a16z podcast. I'm Daisy Wolf, partner on a16z's Bio and Health team, joined by Eva Steinman, investor on our Bio and Health team as well. Today, we are talking with Engy Ziedan, co-founder and chief scientific officer of Protege, and a healthcare economist, assistant professor at Indiana University, whose work has been featured in the New York Times and cited by the CDC.
Today, we are going to dig into why medical AI has a measurement problem, why acing benchmarks does not make a model ready for the hospital, and how Protege is building the referee. Engy, welcome to the podcast.

**Engy Ziedan** (2:17)
Thanks for having me.

**Eva Steinman** (2:18)
Engy, let's start with your background. How did you first meet the Protege team? What were you doing at the time, and how has it evolved since then? Yeah.

**Engy Ziedan** (2:25)
I met Bobby when I was two years out of my PhD, an assistant professor at Tulane in my lonely office, and a pandemic had just hit. I decided that I was going to write papers really fast. So I wanted data basically from two days ago, and that was impossible to obtain at the time. In his past life, he used to work for a data facilitation company, and he offered AWS instances and access to data and connections, and it was a lovely experience. Then we never spoke again, maybe after 2021, 2022, and suddenly in 2024, February, I get this e-mail sent from a Gmail from Bobby, and it has an idea, a Google document in it. In typical Bobby style, it was human written, very simple, and it said something like, what do we hold to be true about the future?
In it, it basically lays out this hypothesis, which has been partly true. When I look back at that document, I often read it every few months or so, which is that models are going to be inhibited in their usefulness by the training data available for them, and that goes beyond healthcare in any domain. It became our mission to provide safe and aligned data, essentially the teachings that would make AI useful for humans.
Today, it's really a proud moment for us. Almost all the models have been pre- and mid-trained on our healthcare data. We provide data in multiple verticals, audio, video, robotics. I go through Slack channels and I see what datasets people are talking about inside Perugia, and it could be anything from 100,000 endoscopy videos inside someone's body to pet data, and I don't mean nuclear medicine, I mean literally veterinary care data to 3D objects, where the team is trying to predict the weight of the object being lifted, so the robot is trained on a diverse set of mugs. It's a fascinating world. Totally.

**Daisy Wolf** (4:20)
By Bobby's background in data, he had the insight very early that the Internet was going to be scraped, and that the frontier models were going to be defined, the best ones would have access to the best real-world data, and a very clear vision of how to get that to people.

27 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID