How Agents Searched 3 Million Epstein Documents | Dylan Freedman
MTS
October 3, 2026
New York Times AI Projects Editor Dylan Freedman breaks down how investigative journalists built an agentic search engine over 3 million pages of released documents to break dozen-plus stories on the Jeffrey Epstein files.
Speakers Dylan Freedman, Theo
TopicsNews
Dylan Freedman (0:00)
When the Epstein files dropped, I mean, the latest batch in the end of January, it was like three million pages of documents. And if you stack that up, it would be like as tall as the Empire State Building.
And so we built like our own Epstein files engine, which is basically an agent that reporters go in, ask a question, it would run SQL in the backend, search for all these documents, interpret them, and then like cite the results with page level citation. So this enabled us to tell like a dozen plus stories about the Epstein files. It's really useful.
Theo (0:34)
All right, we are back. We are live with Dylan Freedman, who is the AI Projects Editor at the New York Times, which is a role that combines reporting with machine learning engineering. Sounds kind of awesome. Sounds sort of vaguely like what I do, sort of. I mean, I don't do much actual machine learning, but you know, I make software with Claude.
And you guys just co-authored this paper, Epstein Files Engine, a genetic search for investigative journalism, which is really interesting. Dylan, welcome to MTS.
Dylan Freedman (1:05)
Welcome. Thanks for having me, Theo. It's good to be here.
Theo (1:08)
Absolutely. So tell us about the Epstein Files Engine. What exactly did you do? What is it? What did you find?
Dylan Freedman (1:15)
Yeah. So basically when the Epstein Files dropped, I mean, the latest batch in the end of January, it was like three million pages of documents. And if you stack that up, it would be like as tall as the Empire State Building. And our job as journalists is to sort of cover it and find what's breaking and newsworthy as quickly as possible. And so this was sort of the perfect use case where AI is actually really useful in an investigative context, you know.
And what we did, very similar to like initiatives like J-Mail, we basically like scraped all of the data.
We collected it, we transcribed it, we used AI to pull out entities, classify emails, things like that. But then we went a step further.
This was like right around the time that agents were starting to get like pretty good. And so we built like our own Epstein files engine, which is basically an agent that reporters could go in, ask a question. It would run SQL in the back end, search for all these documents, interpret them, and then like cite the results with page level citation. So this enabled us to tell like a dozen plus stories about the Epstein files. It's really useful.
Theo (2:24)
Interesting. Like what in particular were you able to find from the Epstein files engine that would have been much harder to do combing through everything yourself?
Dylan Freedman (2:36)
Yeah. I mean, this type of thing, like there were over 20 reporters on it just full time and they all sort of got tasked different people to investigate.
A lot of these people, like you can look them up in the, just like a keyword search through all the documents, and you're not going to necessarily find them. Just because the OCR is kind of weird, it might be like in images. And sort of AI enabled us to make connections that were harder otherwise. We did like a semantic index over all the text and over all the images and the files and things like that. We built sort of all this tooling on top of it.
Within like a week, it was super vibe coded. But in a way that was still useful and tracked like the citations as we went so everything could be traced back and the reporters could verify everything that they found.
Theo (3:24)
What other things do you use AI for at the New York Times? Like you personally?
Dylan Freedman (3:30)
Yeah, so my beat is sort of more on the reporting side. Like the AI initiatives team that I'm on were sort of split into a bunch of things. When people work on tools to help edit, like add alt text, survey journalists on how they're using AI or design products. I've been mostly partnering with reporters based in Washington DC. So a lot of political stories.
Just seeing what they're going through, kind of getting boots on the ground experience, and then being like, here's the ways AI could help. So most of the things I look at are like back-end research focused, like sifting through millions of documents, images, multimedia super quickly, being able to surface insights that are relevant, that sort of thing.
Theo (4:14)
Right. Yeah, I mean, we find this stuff super helpful too at MTS. Like, I basically have like a daily brief that I get every morning from various different agents. Right now, I'm using Opening Eyes Dot for it, but previously it was Grokbot and Claude. And I basically have it like search through all of the new sites and pull together the top stories. And it does a really good job, which is not the case even three months ago ish, two, three months ago.
15 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Fetch the whole transcript
The demo key returns a sample episode in full, no card needed:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090Markdown with the speakers named, for your notes, your knowledge base, or anything that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000793016694