**Erik Torenberg** (0:01)
Hey, everyone, Eric here. We're really excited about a new AI show from Turpentine called Autopilot, hosted by Will Summerlin.
This podcast explores the adoption and rollout of AI in the industries that drive the economy, and the dynamic tech founders bringing rapid scalable change to slow moving industries, from law, to hardware, to aviation. Will interviews founders backed by Benchmark, Greylock, YC, and more to learn how they're automating at the frontiers in entrenched industries. Click on the link in the description to subscribe to Autopilot.
**Jungwon Byun** (0:31)
We have big visions of transforming research, and there might not be a lot of time left before we really need to make it useful for very high stakes decisions and times of chaos.
**Andreas Stuhlmüller** (0:40)
People probably still are making the obvious mistake of being a K model, say yes or no, and then justify your answer. And that's obviously terrible because then it has to answer without any reasoning, and will just lock itself into a potentially wrong avenue.
**Jungwon Byun** (0:52)
This is where the task decomposition comes into play, because it's much easier to iterate on and ship a new predefined column for statistical technique use than it is to ship a new model that's been fine-tuned on all those scientific papers and evaluate that for quality. So again, breaking down things into small tasks helps with launches as well.
**Andreas Stuhlmüller** (1:11)
The key question is, how do we as a society want to turn compute into more work? And one answer is we're going to train larger and larger models, and we'll hope for the best. Maybe we'll augment it a little bit with better interpretability methods.
We are working towards a different answer which is we think like more compute will in fact lead to great things and more correct answers, et cetera. But we can get there through more transparent architectures if we can build the right infrastructure.
**Nathan Labenz** (1:36)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Eric Thornburg. Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to present a conversation with Andreas Stuhlmüller and Jungwon Byun, co-founders of Elicit, a company that helps social and natural scientists analyze research papers at superhuman speed and in which I am a small-dollar but very proud investor. Andreas and Jungwon bring deep expertise and a highly principled approach to the challenge of getting large language models to reliably answer complex research questions. Their product allows users to systematically search large bodies of literature, extract key information into well-structured tables, and iteratively refine queries to zero in on the most relevant sources, all while maintaining a transparent log of each step for easy auditing, reproducibility, extensibility, and team collaboration. As the AI research assistance base has become increasingly crowded, Elicit stands out for its meticulous approach to task decomposition. That is the art which they are very much developing into a science of breaking big questions down into smaller, more manageable subtasks that language models can reliably execute. This allows the product to provide real value to serious researchers who really need to be able to trust the results. We cover a lot of ground in this conversation, starting with the company's founding vision of using AI to enable knowledge sharing at an unprecedented scale. Andreas and Jungwon explain how they have approached key challenges like minimizing hallucinations, maximizing accuracy, and expanding the scope of analysis. And we dig into the tech stack that makes this possible, including their approaches to retrieval, extraction, summarization, synthesis, and, most importantly of all, evaluation. Finally, we touch on the company's evolution from nonprofit research lab to mission-driven commercial startup, their plans to enable more scalable, compute-intensive workloads in the future, and what sorts of talent they are looking to hire today. By focusing on researchers who are working on hard, sometimes literally life-and-death problems, Elicit has no choice but to treat reliability as a top priority. And their approach, honed over several years of R&D, is a model for anyone who's looking to build high-stakes applications with large language models today.
I was obviously a big fan of the company coming in, but I came away from this conversation even more convinced that Andreas, Jungwon, and the team are up to this important challenge.
As always, if you find value in this work, please share the episode with your friends. This one would be great for anyone who's struggling to keep up with an exploding body of research literature, whether that's in machine learning as it is for me, biology, or something else. And please also take a moment to share any feedback or topic suggestions that you have via our website, cognitiverevolution.ai, or by messaging me on the social media network of your choice.
77 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000651291485