**Daphne Koller** (0:07)
In some sense, the wealth of opportunities here is one of the biggest challenges, because everywhere you look, there is a big opportunity for Machine Learning to be deployed in a potentially quite significant way. It's like computers, you're going to use it everywhere, and it's going to be transformative everywhere.
It's not going to be the silver bullet unless you figure out how to use it most effectively, but the opportunities are pretty much endless.
**Sarah Guo** (0:36)
This is the No Priors podcast. I'm Sarah Guo.
**Elad Gil** (0:39)
I'm Elad Gil.
**Sarah Guo** (0:40)
We invest in, advise and help start technology companies.
**Elad Gil** (0:43)
In this podcast, we're talking with the leading founders and researchers in AI about the biggest questions.
**Sarah Guo** (0:55)
We've talked about computational biology for decades, but drugs keep getting more expensive to discover. And at the same time, recent advances in using machine learning for the life sciences and medicine are extraordinary.
Are we on the verge of a paradigm shift in biotech? We're thrilled to have a pioneer in AI, Daphne Koller, on the show to help us explore that question. She's CEO and founder of Insitro, a company that applies machine learning to pharmaceutical discovery and development, specifically by leveraging induced blurry potent stem cells, which we'll get into explaining.
Daphne was a computer science professor at Stanford, co-founder and CEO of Coursera, is a MacArthur fellow, and was named by Time as one of the world's 100 most influential people. We could go through all her other work, but we'd run out of time. Daphne, welcome to the podcast.
**Daphne Koller** (1:37)
Thank you, Sarah. It's a pleasure to be here.
**Sarah Guo** (1:39)
As we were saying, we won't ask you to walk through every part of your amazing life story, but you came to biology as a computer science application years into your career. What sparked you going down that route?
**Daphne Koller** (1:49)
My initial interest in biology came from the technical side in the sense that the data sets, this is way back when, in the mid-90s, the data sets that were available to machine learning research at the time were kind of boring and not very inspiring. So things like classifying text into 20 different newsgroups.
And I found that there were more interesting data sets technically to be had on the biology side back then as we were starting, for example, to measure the activity of genes across the entire genome in multiple samples. So initially, it was really more from a technological perspective, but then I ended up actually having an interest in biology in its own right and ultimately ended up having a bifurcated lab at Stanford where half my lab did core machine learning work published in traditional computer science venues and the other half did core biology work that was published in Nature and Cell and Science. And what was really interesting is that most of my computer science colleagues had no idea that I did biology. Most of my life science colleagues had no idea I was in a computer science department. So it was a bit of a bifurcated existence, but it was a lot of fun.
**Sarah Guo** (2:53)
One more historical question for you. You wrote the book on probabilistic graph models.
When I asked a mutual friend what I should ask you, he suggested what motivated that work and how that field has changed.
**Daphne Koller** (3:04)
Just like in most fields, there is a swing of a pendulum. A lot of the early work in probabilistic graphical models was hugely influential in bringing artificial intelligence more into the world of machine learning and working with numerical data rather than just symbolic AI.
Then I think the advent of deep learning pushed that to the side a little bit because there was so much power that could be gained from basically the pattern recognition from raw inputs, raw images, text and so on without having to worry very much about interpretable representations. What I think we're starting to see right now is a pendulum starting to swing back in the sense that there is a greater understanding that you really need a bit of both. You need that hugely powerful pattern recognition that we get from deep learning, but you also need the ability to reason about things like causality, and you also need some interpretability of your deep learning models so that you can potentially convey to a clinician why you made the decision that you did. What we're ending up with as a really powerful paradigm is some kind of synthesis of the ideas for both of these disciplines coming together.
**Elad Gil** (4:14)
You went from Stanford, I believe to then going and co-founding Coursera with Andrew Ng, and then you went to Calico a few years after that. I'm curious, what made you decide to go into Calico because you mentioned your career was split between life sciences and computer sciences, and so you went on the computer science online learning route, and then you went back into biology, so I'm a little bit curious what drove you back in.
39 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000602459686