Why Universities Must Evaluate Frontier AI | Chris Manning
MTS
September 17, 2026
Sophia Dew interviews Stanford NLP Group founder Chris Manning on why universities are uniquely suited to be independent AI safety evaluators, the risks of client capture with existing IVOs like METR, and how Moonlake is combining code-based physics simulation with neural rendering for physical AI.
Speakers Chris Manning, Sophia Dew
TopicsNews
Chris Manning (0:00)
If we just think of the whole space of Frontier AI, a huge problem at the moment is so much of it is happening in these sort of three monasteries, and very little information is making it outside to the rest of the world. And I mean, I think that's an enormous problem for the whole of society, right? The American public is not gonna get more comfortable with the possibilities of an AI future if the feeling is, it's completely hidden from them.
Sophia Dew (0:33)
Hello everyone, and welcome back to MTS. Today, I am joined by Chris Manning. He founded the Stanford NLP Group and has been one of the major researchers shaping modern language AI for decades. And this week, right after Dario called for outside evaluators inside Frontier Labs, Chris actually proposed that the Stanford NLP would be the right people for the job. And they argued that the important parts of this work, universities would be better than any other organization for actually completing these tasks. He's also a distinguished member of technical staff at Moonlake, where he's working on world models and simulation for physical AI. So Chris, welcome to MTS.
Chris Manning (1:11)
Thank you and great to meet you, Sophia.
Sophia Dew (1:13)
I was so excited to have you. I know we didn't get a chance officially to chat and hang out when we were on campus, but you were pretty famous when you were on campus. You're like the Stanford NLP professor and you're very well known. And so you had some great classes. And so such a great opportunity that we actually get to chat today.
Chris Manning (1:33)
Yeah, absolutely. And I recommend to everybody, if you can't make it on the Stanford campus, you can see CS224, Natural Language Processing with Deep Learning, either on YouTube or by taking as a professional class, often by Stanford Online.
Sophia Dew (1:48)
No, you'll have to. CS224. I'm jealous. Now, all these classes are officially on YouTube. People outside, you don't even need the degree. You don't even need to be a student there. You can actually take them. So and I definitely recommend for anyone watching this, take these classes because they'll be even more relevant now than ever before.
But yeah, I'd love to start a little bit about what you shared the other day around proposing that Stanford NLP could be this independent third party evaluator and kind of goes up with exactly what Dario said was needed. And you even listed out the specific reasons why universities are better than other types of organizations for this work. Can you summarize a little bit of your core argument here?
Chris Manning (2:26)
Sure. I mean, let me scope it first. I mean, what in the post, I differentiated two parts to the role since for the evaluator, part of what was suggested was a monitoring function where you were checking that the company was satisfying all the things that was meant to be doing and saying it was doing for testing models. I mean, universities just aren't set up to do a monitoring function. That's not the kind of thing we'd be good at. But where I thought that there was a really good role for us and a better fit than some of the other organizations are out there, is for this task of really testing models to work out where there were problems in their alignment, where there were risks in the kind of functionality they had. And a lot of that comes down to the fact that that's in the wheelhouse of what university research is about, that we're trying to poke at things and find out what's wrong for it, and where there are opportunities to improve things. And I believe that a university could just do that with more creativity and more rigor than any other organization, and with a fraction more intellectual independence, and that that would be great. Yeah.
Sophia Dew (3:50)
I would love to also get a little bit of your comparison here. So what are universities uniquely good at? How does this maybe compare if Stanford NLP were to be doing, was the IVO as opposed to meter? Where do they actually differ?
Chris Manning (4:04)
So, I mean, there are a number of aspects in which we differ. The one I want to emphasize most is intellectual rigor and creativity.
That I think that university research is just the high point of where people are rigorous and careful and test hypotheses to see if they're really true, and consider alternatives and debate them back and forth, and actually really try and figure out what's going on. And so that's what's going to be increasingly needed as these AI systems not only become very powerful, but also there's an enormous amount of sort of heightened positioning as to what they can and can't do.
21 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Fetch the whole transcript
The demo key returns a sample episode in full, no card needed:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000790270438