**Erik Torenberg** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to share an in-depth conversation with Siyu He, post-doc at Stanford in Biomedical Data Science, and lead author of two notable recent papers that use intricate machine learning systems to shed light on two of the great questions in biology. What causes what at the cellular level? And how do the interactions between individual cells give rise to higher level tissue development and disease? The first paper, Squidiff, allows us to predict how cells will respond to perturbations by conditioning a diffusion model, trained to predict transcriptome states or the level of gene expression that a cell exhibits across more than 30,000 individual genes, on semantic directions derived from other experiments. This approach, based on a very modest amount of data, can save researchers months that would otherwise be needed to culture specific cell types and can also shed light on intermediate states that can be extremely challenging to measure directly. The architecture will be familiar to students of image generation models, and also reminds me quite a bit of the mind's eye models used to reconstruct images from brain scan data that we've covered on previous episodes. This is another data point showing just how general purpose today's architectures really are. How models with quote unquote intuitive understanding of important problem spaces like transcriptomics can perform many orders of magnitude faster than physics based simulations, and ultimately how much future discoveries in biology will be made initially in silico and then confirmed by much slower and more costly wet lab ground truth experiments. The second paper, CORAL, makes it possible to combine low resolution tissue data and high resolution single cell profiles into a single integrated view. This is a particularly intricate system that uses graph neural networks and ultimately deconvolves the lower resolution tissue data into a detailed cell by cell picture, thus enabling closer and more insightful study of the tissue as a whole. This project makes notable use of synthetic data and reminds me of the project we covered a few months back from Michael Levin and his co-authors, which had connected gene, drug, and disease information to make previously unknown connections. Perhaps most telling, CUN team completed both of these projects under the supervision of Professor James Hsiao, who we recently had on the show to talk about the Virtual Lab project, which used a team of AI agents equipped with specialist biology models to develop new nanobodies capable of treating emerging COVID variants, and also the protein language model interpretability project, InterPLM. That all of this work could come from one group in such a short period of time supports the broad theme that everything is working in AI. With just a bit more acceleration, perhaps driven by the inevitable generalizations and scale-ups of the techniques we discussed today, we really could see a century's worth of biology progress in just the next decade. As always, if you're finding value in the show, we'd appreciate it if you'd share it with friends, post about it online, or leave a review on your favorite podcast platform. Of course, we always invite your feedback and suggestions too. You can reach us via our website, cognitiverevolution.ai, and you're always welcome to DM me on your favorite social network. Now, I hope you enjoy this deep dive into cutting-edge architectures that are advancing the frontiers of cell and tissue biology. With Stanford researcher Siyu He, lead author of Squidiff and CORAL.
Siyu He, postdoc at Stanford in Biomedical Data Science and lead author of the recent AI for Biology papers, Squidiff and CORAL. Welcome to The Cognitive Revolution.
**Siyu He** (3:50)
Thank you, Nathan. It's great to be here and very excited to be here to share my research to a broader audience. Cool.
**Erik Torenberg** (3:59)
Well, we've got a lot to unpack, so let's maybe just start with a little bit of big picture context setting. I think the audience, and I can also just speak for myself, I've been obsessed with AI for the last few years, like studying it intensively. I feel like I have a pretty good general understanding of the landscape. On the biology side, first of all, the landscape is even bigger and more complicated, and I haven't spent nearly as much time in it. So I always like to just start off by trying to get a sense for like where you think we are in the big picture and where you think like this work is in the big picture. And I guess I can just tee you up a little bit more by saying the Squidiff paper is about the transcriptome state of the cell operating at like the level of a cell, which is quite interesting. And then the CORAL paper is even zooming out to a little bit bigger scope of analysis than that and looking at like samples of tissue.
80 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000700451757