**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm excited to welcome back Andreas Stuhlmüller and Jungwon Byun, co-founders of Elicit, the AI platform for scientific research that's on a mission to radically improve the quality of reasoning that supports high-stakes decisions.
Elicit was founded on the belief that process supervision, where models are evaluated and rewarded for the quality of their step-by-step reasoning, rather than just their final answer, would improve the consistency, reliability, and legibility of AI workflows. Of course, with the rise of reasoning models, which can do much larger and more challenging tasks, but generally hide their chain of thought from users, Elicit faced a challenge. How to harness the power of frontier models while still ensuring that famously unwieldy LLMs actually do what they're supposed to do?
Their answer, as you'll hear, is an interesting synthesis. By creating a DSL, or a Domain Specific Language, that defines reasoning primitives, which they can then deliver and optimize as discrete reasoning microservices, they allow frontier reasoners to dynamically create structured workflows that are then guaranteed to run as defined.
Today, they work with seven of the top 20 life sciences companies, supporting everything from the ranking of candidate drug targets to the defense of drug launch and pricing decisions for regulators and payers. And now, the frontier is shifting to external world models, structured representations which can take a variety of forms that make a model's understanding of complicated bodies of evidence as explicit and self-consistent as possible, with the goal of supporting reliable causal and counterfactual analysis. As Andreas puts it, these world models are a form of continual learning that humans or other AIs can inspect and understand.
Of course, we cover a lot of important details along the way. From the reasons that they believe that LLMs are still too easy to push around to serve as reliable decision support tools on their own, how Elicit thinks about evaluating the source and quality of new and at times contradictory evidence, the promise of certificates of reasoning that would prove that the appropriate reasoning steps were in fact carried out as intended, how Elicit is automating their own work with a system that they call The Line, which now delivers 30 to 50 code changes per week, and their goal of getting this system running well enough that the company continues to make progress during the human's year-end vacation. Plus, how much they're spending on tokens, as a company and individually, where Gemini fits into their stack, and why they are optimistic that legible reasoning will win out over Neuralese in the end.
Andreas and Jungwon are really exceptional at making time to zoom out and consider the big picture, even as they run their company day-to-day. And their hope and bet is that if we prioritize truth-seeking now, we may be able to create a positive feedback loop in which better reasoning begets better reasoning, and that this could still happen in time to steer the singularity in a positive direction. Very few for-profit teams have been as consistent and disciplined in pursuit of their mission, despite the rapidly changing AI landscape. And so, for many reasons, I hope they are right and successful. With that, I hope you enjoy this conversation about how fluid intelligence can orchestrate trusted reasoning workflows, and why continual learning might be best instantiated outside of the model weights. With Andreas Stuhlmüller and Jungwon Byun, co-founders of Elicit. The Cognitive Revolution is brought to you by Mercury, the fintech that more than 300,000 ambitious companies and individuals trust to run their finances. I've wired AI into nearly every corner of my life. My email, my messages, my calendar. I even gave Mercury virtual cards to my agents, with low limits and category and merchant restrictions, for their autonomous use. But still, my AI's access to my financial data has remained limited. With a normal bank, I might export a bunch of statements and have my assistant process them for me. But for real-time, up-to-date information, and certainly for taking any action, trying to get your agent to use the bank via the browser is just too hard, too slow, and too error-prone to be worth it. And that's why Mercury's new conversational interface, Command, is such a big deal. It's built directly into Mercury, which means you get natural language access to your finances without exposing anything outside of your bank account. No exports, no spreadsheets, no pasting your transactions into third-party tools.
I really think a lot of people are going to prefer it this way. And it can already help you take actions too, with everything bound by the permissions and approval policies that you've already set up in your account. I am genuinely impressed to see this level of AI integration in banking in 2026 And so I invite you to join me in the future. Visit mercury.com to learn more and apply online in minutes. Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column NA, members FDIC. Thank you to Mercury for supporting the Cognitive Revolution. And now, on with the show.
91 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000773175145