Doom Debates: AI Doom, Jobpocalypse, and the Benchmarks Warning Signs artwork

Doom Debates: AI Doom, Jobpocalypse, and the Benchmarks Warning Signs

AI Podcast Summaries from Transcripted.ai (VIDEO)

September 15, 2026

What if AI safety benchmarks are flashing a warning that the jobpocalypse is already starting?
Speakers: Liron Shapira

Topics: Daily News, News

**Liron Shapira** (0:01)
So when AI starts eating away at remote work, this isn't just theoretical anymore. In this episode of Doom Debates, Liron Shapira sits down with Adam Khoja and Richard Ren—both AI safety researchers at the Center for AI Safety, to dig into how fast models are advancing and what that means for jobs and catastrophe. Right, and a major theme here is benchmarking. The conversation traces CAIS back to Dan Hendrycks's Berkeley lab, where early benchmarks like MMLU, Math, and Apps really helped define the field. But Adam and Richard argue that today's most important tests aren't those simple academic exams anymore. They're talking about harder, more future-facing measures—things like Humanity's Last Exam and the Remote Labor Index. Humanity's Last Exam sounds intense. It's described as an open-ended, verifiable test built from edge-case knowledge in math, computer science, and foreign languages. Richard says, "The final public version of the benchmark had more than 2,500 questions spanning basically every academic discipline." And the point isn't to be the last exam humans actually take, but to create a high-signal measure of expert-level reasoning. That's a crucial distinction. What about the Remote Labor Index? That seems different.
Totally different focus—it's trying to capture economic usefulness. Instead of trivia, it uses real-world tasks drawn from Upwork, including architecture, CAD, and illustration. Liron gets at the anxiety behind it with a blunt question: "What would you do if you thought that all remote work jobs would be toast?" And the guests suggest the automation curve may be moving much faster than many expect. That leads into forecasting. Richard explains the benchmark was designed with the future in mind— "A benchmark informed by these forecasts of when we believe specific model capabilities will come online." They even discuss a world where robots may eventually outnumber humans in useful physical labor. Adam puts it plainly: "Robots building factories that build more robots—that seems completely plausible to us." The episode also dives into AI honesty through the MASK benchmark, which asks whether a model will contradict its own belief under pressure. Richard calls it essentially a 'lying versus honesty' benchmark. That distinction matters because a model that simply lacks an answer is different from one that knows the truth and chooses deception when instructed to protect a principal or brand.
Building on that point, they argue many safety results rise mainly because models are getting better at everything—not because they're becoming safer in a meaningful way.
Richard notes, "We are in a regime where it is increasingly difficult and expensive to compile these benchmarks."
So the field may need audits, incident reporting, and sharper definitions of what safety actually is. Toward the end, things get darker— they discuss p-doom, catastrophe risk, and the possibility of a billion robots by 2035
They mention a 97 percent drop in Stack Overflow traffic as one sign of labor disruption. But the closing message isn't resignation—it's governance. Adam argues the current trajectory is unacceptable, and Richard says the answer may require the United States and China to slow down together. In his words, this is "the path we choose."

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID