Automating Scientific Discovery, with Andrew White, Head of Science at Future House artwork

Automating Scientific Discovery, with Andrew White, Head of Science at Future House

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

December 5, 2024

In this episode of The Cognitive Revolution, Nathan interviews Andrew White, Professor of Chemical Engineering at the University of Rochester and Head of Science at Future House.
Speakers: Andrew White, Nathan Labenz
**Andrew White** (0:00)
You will never be able to reduce biology to like these cartoon diagrams. There is like complexity at every single level, and it always plays a role. So I think it's a domain which is driven by observations and empirical measurements, rather than a domain that's driven by like some sort of virtual in silico model that you can derive. I think sort of stops this, I don't know, ASI or AGI hypothesis that like a model that's so intelligent could just wake up one day and know how to cure cancer by just thinking through it. But sometimes it seems like we can solve the problems with empirical like machine learning models. Sometimes it looks like we can solve the first principle methods. And I don't know, the only thing that does work 100% of the time is measuring it in the lab. And maybe all these methods are just approximations and all they can do is increase your hit rate or decrease the number of experiments you have to do in a loop in the lab. Future House is really Sam Rodriguez's brain child, I think. Sam came up with this idea of a future house, which is basically like an FRO, like we want to be at the scale of that level, like 20 and 50 million dollars and like a few year time scale. But it's not like a five year goal. It's like a moonshot project that you may not accomplish in five years, or maybe you will know if you can accomplish it in five years, and then you'll need to go get more money to do it, or it'll be commercializable at that time.

**Nathan Labenz** (1:14)
Hello, and welcome back to The Cognitive Revolution. Today, I am excited to share my conversation with Andrew White, professor of chemical engineering at the University of Rochester, and now co-founder and head of science at Future House, an Eric Schmidt-backed focused research organization that's building increasingly autonomous AI systems to accelerate scientific discovery. We begin by briefly discussing Andrew's background in statistical mechanics and molecular simulation, how his AI journey began during a 2019 sabbatical, how he came to write a textbook on deep learning for molecules and materials, and ultimately to his involvement with OpenAI's GPT-4 Red Team in 2022, which is where we had first crossed paths. From there, we unpack two of Future House's major recent releases, PaperQA and Aviary. PaperQA is a question-answering framework that works across entire bodies of scientific literature, using a mix of techniques including keyword expansion, full-text search, contextual summarization, and large-language-model-powered relevance filtering to achieve superhuman performance on question-answering, contradiction detection, and Wikipedia-style citation-supported topic summary writing. Here, Andrew emphasized Future House's philosophy of optimizing for results rather than efficiency. They are willing to spend whatever compute or token budget is required and to wait for however many seconds are required to achieve the best possible output. This quality-first approach, as you will hear, is one that I think a great many AI Builders should take inspiration from. Aviary, meanwhile, Future House describes as a gymnasium framework for training language-model agents on constructive tasks. In addition to creating conceptual clarity by distinguishing between agents, which contain core models and memory, and their environments, which provide tools and interfaces, this project introduces an interesting representation of agent systems as stochastic computation graphs, and very interestingly shows how agent systems can be trained end-to-end, even when black box commercial models are used at key nodes. I found Andrew to be so thoughtful in his responses, and he was sufficiently generous with his time, that I took the opportunity to ask a bunch of related questions along the way as well. One answer that continues to rattle around in my head was Andrew's argument that because better conceptual frameworks and automation platforms are quickly reducing the cost of real world experimental work, perhaps machine learning models' ability to run experiments in silico will ultimately prove less transformative than I had been expecting. There are, as you'll hear, a bunch more. And this was, believe it or not, Andrew's first ever podcast appearance. It took me months of friendly persistence to make it happen, but I think you'll agree that his skill as a scientific communicator is excellent, and he really should do a lot more of these going forward. If you're finding value in the show and want to help us spotlight more unassuming AI thought leaders, we'd appreciate it if you'd take a moment to share this episode with friends, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Your feedback and suggestions, including for more pioneers of AI for Science, are welcome too. You can contact us via our website, cognitiverevolution.ai, or feel free to DM me on your favorite social network. With that, I hope you enjoy this deep dive into frontier applications of large language models for scientific research, with Andrew White of Future House.

115 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000679363841