$75k Contest Launch! ChinaTalk Hiring + Evals for the Situation Room artwork

$75k Contest Launch! ChinaTalk Hiring + Evals for the Situation Room

ChinaTalk

August 10, 2026

Contest links: General $50k https://www.chinatalk.media/p/50k-chinatalk-submission-hiring-contest we're looking at submissions on a rolling basis! AI evals: Due Over the past few years, we’ve seen hints of policymakers and national leaders using AI models in their actual policy decision-making.
Speakers: Jordan Schneider, Florian Brand, John Chen

Topics: Politics, News, Technology

**Jordan Schneider** (0:00)
Exciting news here at ChinaTalk. Thanks to a generous new grant from Coefficient Giving, ChinaTalk is now hiring and launching two contests to celebrate. The first is a $50,000 contest to spark new research and find talent to work here. And the second, 25 grand for AI evals for the Situation Room. On the first one, we want you to submit work you think will impress the ChinaTalk team. We're interested in proposals related to AI and China, broadly speaking, which include models, chips, robotics, supply chains, applications, as well as AI and the politics, economics, and governance of the technology. Content is format agnostic. We are excited to see you are writing websites, trackers, policy proposals, new datasets, videos. By the way, we're hiring especially for a video producer to do both shorts as well as mid-length YouTube docs, or just surprise us, come up with something cool. For what it's worth, we have a soft spot for graphic novels and games. So while this contest is basically anything you think the team at ChinaTalk would love, the second one on AI evals for the Situation Room requires a little bit more explanation. So we recorded a whole show with two AI evals experts to get you excited about that one.
So over the past few years, we've seen little hints of policymakers and national leaders starting to use AI models in their actual policy decision making. The Prime Minister of Sweden said he used it for second opinions on policy. German Chancellor also said the same thing. Even Trump said he had it like write a speech for him at some point. I think it is fair to assume that senior leadership across the world have started to use AI, not just for tactical things, not just for operational things, but increasingly for broad strategic decision making. While that is exciting and there is a promise of uplift and smarter calls being made on some of the most important decisions that leaders have to face with regards to foreign policy and national security, we're also flying blind. There is an enormous amount of effort and energy that goes into benchmarking and evaluation for stuff like coding.
Also the experiments that you can run when trying to improve your models to get better at coding and software development are much easier to execute, have much lower stakes than running a real experiment when you're thinking about invading a country or signing a treaty or whatnot. Which is why we are at ChinaTalk, trying to kickstart a field which is aimed at better allowing researchers and policymakers to understand just exactly what it is they're working with when they ask these models to support them with some of the most consequential decisions they may make in their lifetimes. So we're going to be launching an essay or I guess an evals slash essay project contest to sort of explore this theme. And over the course of this hour here, I brought on two expert AI eval creators to sort of discuss their views on why the field is important, what interesting work has already been done to start to get at these types of questions of how models approach these broad national security strategic questions, and how you, either as a eval professional, semi-professional, or just concerned person, can start to contribute to this field and come up with new ways to poke and prod at these models and see what they can really do. So joining us today, we have Florian Brand, research engineer at Prime Intellect, as well as John Chen, professor at the University of Arizona, who's done some pretty wild things, getting models to start nuclear wars with each other in Civ V. Let's be clear.
Welcome to ChinaTalk YouTube. Florian, what is an AI eval and why do they matter?

**Florian Brand** (4:16)
So AI evals try to put some number at something we want to measure. That can be things like knowledge questions, which we can measure with multiple choice, or that is the current state where we are looking into agents, where we want to see how good Claude code is at implementing rather complicated code bases these days. So over the last few years, evaluations have gotten more professional and more complicated to reflect the reality that the users are using the products for.

**Jordan Schneider** (4:54)
So let's explore that, right? Because we started out in a world where you would make up these hard science or math problems, and slowly but surely, the models got better at them. The problem, I guess, is once we're getting at these sort of questions where there isn't a right answer, it's not just necessarily that you're trying to see how many points out of a hundred a model would spore, but evaluating a model almost as you would evaluate a hire or a person, there are much softer things you're trying to get at nowadays, right?

32 more minutes of transcript below

Thousands of transcripts fetched by people building searchable podcast archives

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID