**James Zou** (0:00)
One nice thing about these virtual agents is that their meetings are much faster than our meetings. In the time that we sit around and introduce each other and have coffee, they already had hundreds of meetings. These 14 language models, I think they have learned new concepts. They have filled in some of the gaps in our knowledge, and now it's our job to see, can we extract those out using techniques like SAEs. I do think that there is already a goldmine of new concepts and knowledge that are already hidden in existing models that even if we can extract those out, I think that would already teach us a huge amount of interesting new insights.
**Nathan Labenz** (0:35)
Hello, and welcome back to The Cognitive Revolution. Before getting started today, I want to take a moment to invite all of you to submit questions for an Ask Me Anything episode that we'll be producing in the next couple of weeks. From practical application to galaxy brain philosophy to parenting and career choices in the age of AI, I may not have all the answers, but it's all fair game to ask. And while you're there, we'd appreciate it if you'd complete a short survey so that we can learn more about the audience and how we can serve you better. All of the questions are optional. It's fine if you want to be anonymous. And we do have a few thank you gifts planned as a token of our appreciation. Today, my guest is James Zou, Professor of Biomedical Data Science at Stanford and Investigator at the Chan Zuckerberg Initiative, who's recently published several Frontier Advancing Papers at the Intersection of AI and Biology. Our first topic today is the Virtual Lab, a framework that uses minimal human oversight to guide an AI research team, comprised of an AI professor, multiple AI specialist agents, and an AI critic through the process of understanding and exploring open-ended research problems. In a truly impressive demonstration of capability, when challenged to develop new treatments for emerging strains of the COVID virus, the system made the somewhat unorthodox choice of pursuing nanobodies, rather than the more commonly used antibodies, devised a novel workflow combining protein language models, alpha fold, and physics-based tools, and ultimately designed more than 90 candidate antibodies, of which two have since been proven by physical experiment to be of high therapeutic potential, due to their ability to bind with new virus variants, while also maintaining efficacy against earlier versions. While the AI here is not entirely autonomous, the human contributor wrote a bit more than 1% of the total tokens for this project. This is nevertheless one of the most impressive demonstrations of large language models actually doing science that I've seen, and one that I take as a sign of much more soon to come.
The second paper, InterPLM, Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders, is actually equally remarkable. By training sparse autoencoders, which we've covered in depth in previous episodes, including most recently with the founders of mechanistic interpretability startup GoodFire, but in this case for protein language models, which are trained purely on amino acid sequences, the team was able to demonstrate not only that mechanistic interpretability techniques such as auto labeling can generalize across modalities, but also that mechanistic interpretability can deliver totally new discoveries. Most remarkably, they were able to identify features corresponding to at least one entirely new protein motif which had not been previously documented in the literature. This to my knowledge is perhaps the clearest demonstration yet that today's machine learning architectures are capable of learning important natural concepts that humans don't know from raw data. It's hard to overstate just how important this capability is likely to be, and I find the implications for how we should understand more familiar large language models to also be extremely profound. Now I only had an hour with Professor Zou, so this is one of our shorter episodes, but the results here are second to none, and I've also recently discovered another paper on which he was a co-author about modeling the evolution of cellular transcriptome states over time, which I hope to cover in another episode soon. If you'd like to help us make that happen, we'd appreciate it if you'd take a moment to share this episode with friends, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback and suggestions, including for more topics at the intersection of AI and biology, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. For now, I hope you enjoy this conversation about truly remarkable recent results in AI for Biology with Professor James Zou. James Zou, Professor of Biomedical Data Science at Stanford and Investigator at the Chan Zuckerberg Initiative. Welcome to The Cognitive Revolution.
53 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000680829979