⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data artwork

⚡️The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals & Human Data

Latent Space: The AI Engineer Podcast

February 23, 2026

Olivia Watkins (Frontier Evals team) and Mia Glaese (VP of Research at OpenAI, leading the Codex, human data, and alignment teams) discuss a new blog post (https://openai.
Speakers: Mia Glaese, Olivia Watkins
**SPEAKER_1** (0:04)
Okay, hi, we're here in the OpenAI studio with Mia and Olivia from the Frontier Evals team, or however you want to introduce yourself. Maybe you want to introduce, name what you do at OpenAI, and we can get it started.

**Mia Glaese** (0:16)
Sure.

**Olivia Watkins** (0:17)
Hi, I'm Olivia. I'm on the Frontier Evals team.

**Mia Glaese** (0:20)
Hi, I'm Mia. I am a VP of Research at OpenAI, and my teams are the Codex team, the Scheme Data team, and the Alignment team, and we work a lot with the Olivia's team on Frontier Evals.

**SPEAKER_1** (0:33)
Yeah, very exciting. By my understanding, you were part of the original team that worked on SWE-Bench Verified as well.

**Mia Glaese** (0:39)
Yeah. Olivia's team, the Frontier Evals team, and the Scheme Data team collaborated on creating SWE-Bench Verified.

**SPEAKER_1** (0:45)
You've seen the evolution of coding benchmarks over time, and I think it was roundabout to the mid to late 2024, we first covered SWE-Bench Verified. These have evolved a lot since then. What's the blog post that you have worked on that we're releasing today?
What is the main thesis that you're pushing out?

**Olivia Watkins** (1:04)
So the main thesis is that SWE-Bench Verified has been one of the North Star coding benchmarks that the field has looked at to measure coding progress. But recently, we've seen that progress has kind of stalled, and basically we realize that this is because the eval is effectively saturated and also highly contaminated. So at this point, we think that it's not really measuring coding performance improvements well anymore, and we think that the field should move away from this towards other benchmarks.

**SPEAKER_1** (1:28)
Like SWE-Bench Pro.

**Olivia Watkins** (1:29)
Like SWE-Bench Pro, yeah.

**SPEAKER_1** (1:30)
Amazing. Yeah, one of the jokes I always have is like there's a group chat with all the labs, and everyone just turns to increment like 0.1 on tracts. And then it's like, OK, well, you have the best coding model, I guess, because you're 0.1% higher. But it's not super convincing at this point. So cool, I think, let's sort of reset on like, what was the original work that you guys did for SWE-Bench Verified, which I think was pretty substantial. Like it was like a very significant investment from OpenAI, which people still don't appreciate. And then what were the dissatisfactions that we found over time, right? So like, what was SWE-Bench Verified that people should know about?

**Olivia Watkins** (2:08)
SWE-Bench Verified was kind of a cleanup of original Bench, academic benchmark from the lab at Princeton called SWE-Bench. And the agent is basically given a code base and a task that was sourced from a real world repository and GitHub issue, and was asked to solve the task and is graded on whether some tests pass. And at the time this was quickly became a popular benchmark because at the time the field didn't really have good real world coding benchmarks. But then when OpenAI took a look at the benchmark as part of one of the evals we wanted to track in our preparedness framework, folks started realizing that some of the cases where agents were failing were due to bad problem setups rather than just to models being dumb. So folks at OpenAI did a pretty extensive human data campaign hiring like almost 100 real world software engineers to go through the problems and figure out like are the tasks well specified, are the tests actually fair, and kind of created a curated set of like 500 tasks that we thought were much better.

**Mia Glaese** (3:06)
It's just maybe it's hard to overstate like the amount of effort that it took to like create that benchmark. It was literally like many expert software engineers reviewing the problems like differentially, multiple times, and to you know, basically like three different experts independently decide of it.

**SPEAKER_1** (3:29)
Yeah, you didn't have to do that. You just tripled your costs for just...

**Mia Glaese** (3:32)
I mean, we had to do it, actually, because it's quite a hard task to like look at something like a problem and the patch. And then like, it's not just the problem and the patch, right? You have to like understand it in the context of the code base, that the human are the models and to solve the task. So it's a very complex problem and it was definitely needed to have three reviews. And I think like maybe we should have done more, but it was definitely a lot of effort to get there.

**SPEAKER_1** (4:00)
Yeah. And there's more, but people can read the blog post for that. I will note that you guys had a trend in verifying benchmarks because I just recently saw, I think Quyen had a HLE verified for humanities license verified.

21 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000751071625