Can AIs Generate Novel Research Ideas? with lead author Chenglei Si artwork

Can AIs Generate Novel Research Ideas? with lead author Chenglei Si

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

October 23, 2024

In this episode of The Cognitive Revolution, Nathan delves into the fascinating world of AI-generated research ideas with Stanford PhD student Chenglei Si. They discuss a groundbreaking study that pits AI against human researchers in generating novel AI research concepts.
Speakers: Erik Torenberg, Nathan Labenz, Chenglei Si
**Erik Torenberg** (0:01)
Hey, everyone. Erik here. We've got something exciting in the works, and we want you to be the first to know about it. Turpentine, the network behind the show you're listening to right now, is launching a publication, and we're offering early access to our listeners. We'll have our biggest hosts and expert guests writing pieces and leverage our group chats for content inspiration. For an early preview, drop your e-mail at the link in the show notes. You can also head to turpentine.co/exclusivedashaccess. Now, on to the show.

**Nathan Labenz** (0:30)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. As a developer, the journey from concept to production-ready large language model apps is fraught with challenges. Dealing with unpredictable language model outputs, hallucinations, and ballooning API costs can all be blockers to shipping your next AI-powered feature. That's where Advanced RAG comes in. With the new RAG++ course from Weights and biases, you can overcome these hurdles and build reliable production-ready RAG applications. Go beyond proof of concept and learn how to evaluate systematically. Use hybrid search correctly and give your RAG system access to tool calling. Based on 21 months of running a customer support bot in production, industry experts at Weights and biases, Cohere, and Weaviate show you how to get to a deployment-grade RAG application. This offer includes free credits from Cohere to get you started. Make real progress on your large language model development and visit wnb.me.cr to get started with their RAG++ course today. That's wnb.me.cr to get started with their RAG++ course today.
Hello, and welcome back to The Cognitive Revolution. Today, my guest is Chenglei Si, a PhD student at Stanford who's developing ways to use large language models to automate research. He's lead author of a fascinating newspaper that asks the question, can large language models generate novel research ideas? This question has become one of the most important in the entire field, as the ability to generate research ideas that are truly worth pursuing, particularly in the domain of AI, has long been considered a key precursor to recursive self-improvement loops and a possible intelligence explosion. Since the early days of GPT-4, we've seen several notable attempts to create AI research assistance, including projects like Co-Scientist from Gabe Gomez's group at CMU, which we did a full episode on, and more recently, the AI Scientist paper from japanese company Sakana AI as well.
these systems demonstrated capabilities that would have seemed impossible just a couple years ago, including translating natural language instructions to chemistry protocols, and using the Semantic Scholar API to assess research ideas for originality. nevertheless, the question of whether AI systems can produce high-value research ideas, true eureka moments, has remained the subject of fierce debate. Chenglei and his collaborators set out to shine light on this subject with an ambitious study. They asked more than 100 PhD researchers working in AI for new research ideas, incentivizing quality with cash prizes for the best ideas, and then asked Claude to generate new ideas as well. After processing the text in an attempt to create a level playing field for evaluation, they then had expert reviewers rate all of the ideas without knowing their source. The results? The AI-generated ideas scored significantly higher on both novelty and excitement. Now, as with many recent AI results, this paper has become something of a Rorschach test. Those inclined to believe in rapid AI progress see it as a major milestone, while skeptics criticize the methodology, somewhat unfairly in my view, but you can judge for yourself as you listen, and more persuasively in my mind, emphasize that this work which focuses specifically on prompting techniques for language models may not generalize to the harder sciences. Personally, I agree that we cannot confidently project this one result onto other domains, but I find the experimental setup here to be quite well done, the statistically significant results seem credible, and I think Chenglei's individual observations are worth taking very seriously as evidence too. Everyone listening to this show should be familiar with the AI Maxim to look at your data, and nobody has spent more time with the raw outputs than Chenglei has. So when he reports that 9 of his 10 favorite ideas from this entire project turned out to be AI generated, and that the AI ideas are generally more out of the box and less grounded in existing work than human ideas, I think we would do well to listen. Overall, my feeling right now is that we're at a sort of tipping point, where the Cloud 3.5 Sonnet and GPT-40 class of models can sometimes, with great effort put into system design and many millions of tokens to earn, sometimes generate meaningful novel research ideas, but not yet in a way that makes frontier research dramatically more accessible or scalable. The next generation of models, starting with the O1 series, seems to me pretty likely to change that. I've been coding with O1 and Cursor a lot lately, and I've been really struck by how effectively O1 can review my entire codebase and my plan for a new feature, and return both a critique of my approach and a recommended alternative that is usually genuinely better. As a community, we're still figuring out exactly what the limits to these capabilities are. But if I had to bet, I'd say that systems like Chenglei's, powered by O1 and a bigger inference budget, might well produce undeniably needle-moving results quite soon. As always, if you're finding value in the show, we appreciate it when folks take a moment to share it with friends, write us a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. We welcome your feedback via our website, cognitiverevolution.ai or by DMing me on your favorite social network. For now, I hope you enjoy this conversation on the emerging reality of AI automated research with Chenglei Si. Chenglei Si, PhD student at Stanford, researching the automation of research. Welcome to The Cognitive Revolution.

71 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000674130830