How Double-Blind Tests Could Secure Frontier AI | Patricia Paskov
MTS
October 5, 2026
Patricia Paskov, Director of Standards at AVERI, discusses the growing need for independent frontier AI auditing, the talent shortage in model evaluation, and pilot double-blind evaluations run with Google DeepMind and Openmind. Turn ideas into software people love.
Speakers Patricia Paskov
TopicsNews
Patricia Paskov (0:00)
We were very excited to put out our first pilot results. This was in partnership with Google DeepMind, ML Commons and Openmind. So what we did is we built on this, but using a Gemini model. So this is like, we used a frontier model, and we also used a real benchmark from ML Commons.
And there are like the proper firewalls in place, such that these are both going into the secure enclave that Openmind makes, and it's privacy preserving, and it kind of gives the guarantees of privacy and security on both sides such that you can run the evaluation and spit out the results.
SPEAKER_2 (0:33)
All right, we are back. We are live with Patricia Paskov, who is the Director of Standards at the AI Verification and Evaluation Research Institute, Averi. We've had Miles Brundage, who's the Director of Averi.
Executive Director of Averi on many times. So it's great to have you on Patricia. Welcome to MTS.
So tell us about your work at Averi. What does Averi do right now, especially after the whole independent verification and auditing vibe shift that we've seen over the last few weeks?
Patricia Paskov (1:03)
Yeah, busy time for Averi. I think we were very excited to have the embedded evaluation announcement come up from Dario and kind of the consensus from lab leaders. So kind of at the core of Averi's work is running pilot frontier AI audits. And we do that a bit instrumentally to inform kind of three streams of outputs that we see, which is policy work, standards work, and then open source tooling.
Standards, I think, are like super critical, especially in the absence of policy. And then as we build up that policy, you know, those standards feed into it. With respect to the open source tooling, I think our hope is to kind of like spur the growing ecosystem of evaluators and give the resources to do things like automated auditing and to just really scale the efficiency and speed of auditing.
And, you know, I think we see ourselves as like running those pilots, but also serving the broader ecosystem as this grows.
SPEAKER_2 (1:55)
How excited are you about automated auditing? Like, how will you be able to trust it?
Patricia Paskov (1:59)
Yeah.
SPEAKER_2 (1:59)
That seems like a big problem.
Patricia Paskov (2:01)
Yeah. I mean, so I was just at a workshop this past weekend on automated auditing, and I think so many interesting questions. And a lot of the questions are like drawing on literature that we've been working with in the past. So I think a lot of LLM as a judge ties in. I think scalable oversight ties into automated auditing tools.
I'm excited about it, like, but cautiously excited. So I think, you know, we should be building tools and also putting human in a loop and measuring how effective the tools are and measuring how aligned the tools are and measuring where they fall short and kind of scaling up appropriately with those guardrails.
SPEAKER_2 (2:33)
Right.
It seems like one of the most common things we hear about independent auditing and verification is like, there's just not nearly enough people in the world doing it. Some people have gone so far as to say, there's like 50 people in the world who are really qualified to do this work, which seems like implausible to me. But do you think this is roughly true? How many people in the world do you think can actually audit from TrueLabs? Okay.
Patricia Paskov (2:57)
Great question. I feel like there are a number of numbers floating around, and it's on Avery's radar to get a better estimate of this. I think though that when people think about auditing, they're largely, and maybe giving this 50 evaluators number, they're largely focusing on the evaluation of the model or the system level for Frontier developers. At Avery, we think of it in three parts. So you can audit the model or the system, you can audit compute, and you can audit governance. And when you think about it holistically like that, there's actually loads of people and skill sets and organizations that can feed into that process that I think are kind of underappreciated right now.
So we do have the Frontier evaluators. Of course, METR does amazing work, Apollo on the bio risk side, SecureBio, Nemesis, and a lot of kind of like underappreciated ones, largely because of NDAs. So a lot of Frontier evaluations to date have been done under NDAs.
So there are more even on the model and system level than are visible. At the same time, there needs to be so much more. So like huge advocate for a lot of growth in that area. But then zooming out to those other layers, I think the skill sets of traditional auditors, like the Big Four, KPMG, EY, there are a lot of skill sets in governance auditing and risk management auditing that we can pull from in the cybersecurity world, tons of skill sets. So I think thinking about it holistically, there is actually a lot of talent out there. We just need to do a better job at funneling it and growing the ecosystem that's specifically dedicated to doing this type of auditing.
16 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Fetch the whole transcript
The demo key returns a sample episode in full, no card needed:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000793252456