**Nathan Labenz** (0:00)
Hello, and welcome back to The Cognitive Revolution. Today, I'm sharing an AI Scouting Report presentation that I recently gave as part of the Law and Artificial Intelligence Certificate Program by LexLab at UC Law, San Francisco. My talk was on day one of the week-long program, and my role was to set the stage for the more focused discussions that followed throughout the week, by giving the most comprehensive and current survey of the AI landscape that I could possibly fit into a single time slot.
If you've seen previous scouting reports, the structure of this talk will be familiar. I again broke things down into the good, including my use of AI to help navigate my son's cancer treatment, the bad, including the rise of deception and other advanced forms of reward hacking, and the weird, including the fact that models now recognize when they're being tested at such a high rate that all of our safety tests are called into question. Before concluding, with a bunch of important questions at the intersection of AI and the law that I personally wish I had answers to, and finally opening things up for Q&A. My goal was to make sure that everyone had an accurate sense of how far AI capabilities have come, both in general and specifically as they're being used in the legal profession, while also highlighting the increasingly hair-raising bad behaviors we continue to see from each new generation of frontier models. I zoomed through 90 slides in just over 45 minutes, and while that might feel a bit overwhelming, the dizzying pace is itself a big part of the point. Even I, as someone who's managed to make it my full-time job to keep up with AI developments, can no longer keep up with everything. And in the course of updating these slides, which I hadn't touched since October just before my son got sick, I was once again amazed by how much has happened in just the last few months. The latest frontier models started to push the frontiers of math and physics, achieved parity with expert professionals on GDPVAL legal and a number of other task types, and started to make general-purpose AI agents really work for the first time. On the other hand, we also got some glimpses of the strange future we are racing toward, with the first public hit piece written by an AI agent about a human, the first explicit public timeline for autonomous AI research from OpenAI, and in the same week, Anthropix retraction of their previous safety commitments and open conflict with the US federal government. One practical tip I learned while doing this is that GROK, if nothing else, is outstanding for Twitter search. Over and over, I asked it to find and link tweets about various topics, and it saved me quite a few hours that I previously would have had to spend hunting and pecking to track down sources. The upshot is that just about every slide contains a link to source material where you can learn more about the eureka moments and bad behaviors in question. Find that link in our show notes if you d like to dig in on anything in particular. But otherwise, buckle up for a breathless overview of the current AI landscape from day one of the Law & Artificial Intelligence Certificate Program by LexLab at UC Law, San Francisco. The Cognitive Revolution is brought to you in part by Google, makers of the Gemini family of models, which have consistently led the industry with their famous 1 million token context window. I've had a number of spine-tingling moments with new AI models over the last few years, but one of the most memorable was using Gemini to vibe-code fine-tuning experiments as part of the Emergent Misalignment Research Project. The codebase was over 400,000 tokens, too much for any other model even to attempt at the time. But with Gemini, I was able to go back and forth, iterating on experiment design and even debugging low-level details with the full codebase in context. Once I realized how effective this could be, all sorts of additional use cases started opening up. For APIs that don't have good AI-friendly documentation, I've scraped full documentation websites with all of their repetition and HTML cruft and had Gemini give me a consolidated but still comprehensive version that fits comfortably into context. With max output length of more than 65,000 tokens, it can usually do this in a single shot. In the context of my son's cancer journey, I've had a single long-running thread with Gemini that even with all the test results that I've uploaded over the last four months is still not even 500,000 tokens. Google has upgraded Gemini at least twice since I started that thread, but in AI Studio, you can upgrade to the latest model at any point along the way. Most recently, I wanted to see if Gemini could help me do a better job hosting this podcast, so I gave Gemini 3.1 Pro the full transcripts of 12 recent episodes and asked it to identify consistent structural weaknesses in my approach that represent opportunity for meaningful improvement. It reported back that I'm prone to monologues, that my longest question was more than 600 words, and that toward the end of those episodes, when I'm running out of time, I often try to ask multiple questions at once, which usually just overwhelms the guest. Whether or not I'll be able to correct those bad habits, you'll have to stay tuned to find out. But in the meantime, think about what long context can do for you, and try Google's latest and greatest model, Gemini 3.1 Pro, in the AI Studio or the Gemini app. Thank you to Google for supporting The Cognitive Revolution, and now on with the show. Thank you for having me. Sorry I couldn't be there in person, but again, I appreciate the kind of intro, and I'm going to try to give you guys the probably fastest talk you've heard in quite some time. I've got 90 slides, and I'm going to try to give you the most comprehensive overview I can of everything that's going on in the AI space, which is to say the least a very tall order. Super quickly about me, I did start this company Waymark, and I now host The Cognitive Revolution. There's some interesting lore around my participation in the GBT4 Red Team, where there's a long podcast and a Twitter thread about that if you want to learn more about the backstory. These days, I also do a little bit of angel investing. My favorite page on the Internet is this case study that my company, Waymark, earned with OpenAI way back in the day when it was still GBT3. We were early adopters of this stuff because at that time, it was really only good for doing simple things like writing marketing copy, but that's exactly what we needed it to do. So we became early adopters and I basically became totally obsessed with the technology as I got to know it better and better. So today, as I said, I'm going to try to do kind of everything everywhere all at once, start with some kind of conceptual stuff and then go into a mix of Eureka moments, bad behavior, WTF moments, and some big open questions at the end. And believe me, there are plenty of open questions. So just briefly on what I do, I call it AI Scouting. And I think because that's the term of my own invention, it does bear a little definition. I would define it as maintaining situational awareness for fun, profit, and the public good. I find it personally extremely interesting. I basically have a never-ending curiosity to learn about this stuff. It has actually worked out to be a somewhat decent business model for me personally. But my real hope is that I can inform others and do my small part to nudge us toward a better AI future by helping other people get calibrated on really where we are in this technology wave because it is coming at us extremely fast. This is just the taxonomy of all the different AI jobs that I've cataloged over time. And don't worry, I will give you all the slides. You don't by any means have to read this. I would say the AI scout role is still one of the more hypothetical or speculative. But we are starting to see CEOs more and more say, hey, I'm hiring a person specifically to keep up with AI developments. And I think once you see all the things on here, you'll see that that's certainly at a minimum, not a crazy thing for some CEOs to be doing.
67 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000755660951