Virtual Cell Models, Tahoe-100 and Data for AI-in-Bio with Vevo Therapeutics and the Arc Institute artwork

Virtual Cell Models, Tahoe-100 and Data for AI-in-Bio with Vevo Therapeutics and the Arc Institute

No Priors: Artificial Intelligence | Technology | Startups

February 25, 2025

On this week’s episode of No Priors, Sarah Guo is joined by leading members of the teams at Vevo Therapeutics and the Arc Institute – Nima Alidoust, CEO/Co-Founder at Vevo Therapeutics; Johnny Yu, CSO/Co-Founder at Vevo Therapeutics; Patrick Hsu, CEO/Co-Founder at Arc Institute; Dave Burke, CTO at...
Speakers: Sarah Guo, Johnny Yu, Nima Alidoust, Patrick Hsu, Dave Burke, Hani Goodarzi
**Sarah Guo** (0:05)
Hi, listeners, welcome back to No Priors. Today, we're here with the CEO, CTO, and core investigator of the Arc Institute, as well as the co-founders of Vevo, to talk about their release of the Tahoe-100, the largest single-cell drug-perturbed dataset ever created, as well as where we are in AI for biology, why we need a virtual cell model and not just protein structure prediction models, and when we should finally expect to see treatments from this growth of use of machine learning in bio.

**Johnny Yu** (0:35)
Hi, I'm Johnny, and I work on single-cell RNA sequencing at Vevo.

**Nima Alidoust** (0:41)
I'm Nima. I'm one of the founders together with Johnny. I'm a quantum chemist by background, but I've converted to being a computational chemist that loves playing with biological data, and we're building Vevo to really do that, to predict how chemicals interact with cells in different biological contexts. Some people call it the virtual cell. That's basically what we're working on.

**Patrick Hsu** (1:04)
I'm Patrick Hsu, one of the founders at the Arc Institute, which is working at the interface of biology and machine learning to try to understand and one day treat complex human diseases, which are most of the major killers.

**Dave Burke** (1:17)
I'm Dave, CTO at Arc Institute, focused on computational biology and building novel AI models for biology.

**Hani Goodarzi** (1:25)
I'm Hani. I'm a co-investigator at Arc. I work very closely with Dave and Patrick to push our virtual cell initiative.

**Sarah Guo** (1:31)
Congratulations, everyone. It's a big day. Let's jump right into it. What is the Tahoe-100 and what is the significance of it?

**Johnny Yu** (1:39)
So Tahoe-100 is the world's biggest single cell RNA sequencing data set and it enables basically a ton of machine learning applications, including things like the virtual cell, but it also enables a lot of drug discovery applications. And broadly, in the context of where I think we are as a field, it's kind of the beginning of a different way of doing drug discovery, of basically understanding how to build medicines and basically bringing AI machine learning people into the mix.

**Nima Alidoust** (2:08)
And maybe something I will add there as well. Over the last 20 years or so, people have accumulated a massive amount of data points when it comes to protein structures, protein function, how drug molecules interact with proteins. One thing that we haven't had as much is how different cells behave in different contexts and how different genes within each of those cells actually functions in the presence of the other genes, in these different biological contexts.
We believe this is the era for that right now. You have seen the emergence of protein language models built on the data cells that have been accumulated over the last two decades. But now is the era for actually having data on cells, how they function, how they interact with drug molecules. And exactly what John is saying, Tahoe is really a landmark data set there that allows us to really measure how drugs interact with different cells from different patient models. And that gives us the ability to build similar models that we built in protein language models, but in the cellular kind of context.

**Dave Burke** (3:07)
If you think about it actually, like in history of AI, it's punctuated by the data sets that come about, right? Like if you think about ImageNet in 2009 that Fei-Fei Li put together, and you look at what that did to drive sort of a nonlinear jump in machine vision, I think the hope here is that by producing data sets, that's particularly perturbational data sets that allow us to elicitate cellular responses, that we'll be able to actually drive forward the ability to model at the cellular level, not just at the protein level. And so I think this is one of those moments, hopefully.

**Patrick Hsu** (3:42)
Yeah, so lots of people have been talking about what those foundational data sets look like for biology, right? And this has been really useful for training protein structure prediction models, like AlphaFold built on CASP, the competition built on top of PDB data. But how do you do this for cells and cellular dynamics, which is really what tells us about biology and how it responds in health and disease? So I think those are the core steps forward where we want to bring up our ability to study higher levels of abstraction in biology, not just the individual molecular machines, but how they operate in the context of an entire cell.

**Sarah Guo** (4:18)
Congrats also to the entire Arc team. Given you are working on both virtual cell models and protein structure prediction, protein language models, can you contextualize a little bit why we need both and where we are in the progress of each?

52 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000695869630