**Erik Torenberg** (0:01)
Hey, everyone. Erik here. We've got something exciting in the works, and we want you to be the first to know about it. Turpentine, the network behind the show you're listening to right now, is launching a publication, and we're offering early access to our listeners. We'll have our biggest hosts and expert guests writing pieces and leverage our group chats for content inspiration. For an early preview, drop your email at the link in the show notes. You can also head to turpentine.co/exclusivedashaccess. Now, onto the show.
**Dan Hendrycks** (0:30)
The idea was that a lot of the linguistic understanding benchmarks were not being sufficiently difficult. Elon's description of it is basically like an undergraduate level knowledge and skill test. There's a benchmark we call Machiavelli, which is largely assessing the propensities of LLM agents. See what sort of decisions do they make along the way? Do they screw people over? Or are they generally nice? Do they lie a lot? There is something going on in making AI systems a lot more reliable to jailbreaking due to specific algorithmic advances that do not necessarily follow from just scaling the model. If they get expert level virologists, I don't know if I want that being released. Sorry to say, because you'll run some substantial risks of bio weapons.
**Nathan Labenz** (1:19)
Hello, and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs, and builders working on the frontier of artificial intelligence. Each week we'll explore their revolutionary ideas, and together we'll build a picture of how AI technology will transform work, life, and society in the coming years. I'm Nathan Labenz, joined by my co-host, Erik Torenberg. Hello and welcome back to The Cognitive Revolution. Today I am thrilled to be joined by Dan Hendrycks, Executive Director of the Center for AI Safety, Advisor to Elon Musk's XAI, and one of the most prolific and influential researchers in AI safety and alignment, Full Stop. While Dan has recently become famous in AI circles for his work developing and advocating for SB 1047, I wanted to use this conversation to highlight what an all-purpose AI powerhouse Dan truly is, and to get his perspective on a number of questions that I've personally been thinking a lot about. And so, this is a long episode in which we cover a tremendous amount of ground, including his early work on activation functions, such as the widely used GELU function, which he developed as an undergrad, and his work on benchmarks, including the insights that allowed him to create in MMLU and math, some of the longest lived and most often cited benchmarks in existence today, as well as what he's up to next with a project called Humanities Last Exam. We also cover his work on AI robustness and alignment, including early work on robustness in image classifiers, which showed how difficult robustness can be to achieve. His 2023 paper introducing representation engineering, a top-down alternative to mechanistic interpretability, which can be used to identify directions in light and space for both monitoring and steering purposes. His 2024 paper with past guests, Andy Zhao and Zico Kolter on circuit breakers, which build refusal behaviors more deeply into LLM weights through a specialized fine-tuning process. And most recently, his work on tamper-resistant training, which aims to make it difficult for users of open-source models to remove refusal behaviors via fine-tuning. And which suggests that it might become possible for companies like Meta to open-source models with durable guardrails built in. We also touch on a number of big-picture issues around AI governance and the geopolitics of AI development, philosophical questions about the nature of intelligence and consciousness, and sociological questions about how AI might help us better forecast future events and generally improve our collective epistemics. We even get a fascinating behind-the-scenes look at how he orchestrated the 2023 AI X-Risk Statement, which was signed by a remarkable set of industry leaders, including the CEOs of DeepMind, Amtropic, and yes, Sam Altman of OpenAI, and which said in full that, ...mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war. I think what struck me most about this conversation was that, despite all his experience and success, Dan's approach remains fundamentally experimental and his outlook very empirical. He places far more weight on data and compute than on algorithmic insights as drivers of capabilities advances, and he remains very open-minded both about how far the current AI paradigm will go and about whether any of the safety approaches we're developing will really work when the chips are down.
Ultimately, as you'll hear, that leads him to believe that we would be well served to adopt a defense-in-depth strategy. As always, if you're finding value in the show, we appreciate it when folks take a moment to share it with friends or write an online review on Apple Podcasts or Spotify, and we welcome your feedback via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. Now, I hope you enjoy this expansive discussion with one of the most impactful figures in AI safety and alignment. This is the great Dan Hendrycks. Dan Hendrycks, Executive Director of the Center for AI Safety. Welcome to The Cognitive Revolution.
140 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000673660402