**Sarah** (0:05)
Hi, listeners, and welcome back to No Priors. Today, I'm with Dan Hendrycks, AI researcher and director of the Center for AI Safety. He's published papers and widely used evals such as MMLU and most recently, Humanity's Last Exam. He's also published Superintelligence Strategy, alongside authors including former Google CEO, Eric Schmidt and Scale founder, Alex Wang. We talk about AI safety and geopolitical implications, analogies to nuclear, compute security and the state of evals. Dan, thanks for doing this.
**Dan Hendrycks** (0:35)
Glad to be here.
**Sarah** (0:36)
How'd you end up working on AI safety?
**Dan Hendrycks** (0:38)
AI was pretty clearly going to be a big deal if one would just think through its conclusion. So early on, it seemed like other people were ignoring it because it was weirder or not that pleasant to think about. It's hard to wrap your head around, but it seemed like the most important thing during this century. So I thought that that would be a good place to develop my career toward. That's why I started on it early on. Then since it'd be such a big deal, we'd need to make sure that we can think about it properly, channel it in a productive direction, and take care of some pale risks which are generally systematically under addressed. So that's why I got into it. It's a big deal and people weren't really doing much about it at the time.
**Sarah** (1:24)
What do you think of as the center's role versus safety efforts within the large labs?
**Dan Hendrycks** (1:30)
Well, there aren't that many safety efforts in the labs even now. I think the labs can just focus on doing some very basic measures to refuse queries related to like help me make a virus and things like that. But I don't think labs have an extremely large role in safety overall or making this go well. They're kind of predetermined to race. They can't really choose not to unless they would no longer be a relevant company in the arena. I think they can reduce like terrorism risks or some like accidents. But beyond that, I don't think they can dramatically change the outcomes in too substantial of a way.
Because a lot of this is geopolitically determined. If companies decide to act very differently, there's the prospect of competing with China, or maybe Russia will become relevant later. As that happens, this constrains their behavior substantially. I've been interested in tackling AI at multiple levels. There's things companies can do to have some very basic anti-terrorism safeguards, which are pretty easy to implement. But there's also the economic effects that will need to be managed well, and companies can't really change how that goes either. It's going to cause mass disruptions to labor and automate a lot of digital labor. If they tinker the design choice or add some different refusal data, it doesn't change that fact. Safety are making AI go well, and the risk management is just much more of a broader prom. It's got some technical aspects, but I think that's a small part of it.
**Sarah** (3:12)
I don't know that the leaders of the labs would say we can do nothing about this, but maybe it's also a question of everybody also has equity in this equation. Maybe it's also a question of semantics. Can you describe how you think of the difference between alignment and safety as you think about it?
**Dan Hendrycks** (3:28)
I'm just using safety as a catch-all for dealing with risks. There are other risks. If you never get really intelligent AI systems, that poses some risks in itself. There's other sorts of risks that are not as necessarily technical like concentration of power. So I view the distinction between alignment and safety as alignment as being a sort of subset of safety. Obviously, you want the value systems of the AIs to be in keeping with or compatible with say the US public for USAIs or for you as an individual. But that doesn't make it necessarily say if you have an AI that's reliably obedient or aligned to you, this doesn't make everything work totally well. China can have AIs that are totally aligned with them. The US can have AIs that are totally aligned with them. You still are going to have a strategic competition between the two.
They're going to need to integrate it in their militaries. They're probably going to need to integrate it really quickly. This competition is going to force them to have a high risk tolerance in the process. So even if the AIs are doing their principles as biddings reliably, this doesn't necessarily make the overall situation perfectly fine. I think it's not just a question of reliability or whether they do what you want. There are other structural pressures that cause this to be riskier like the geopolitics.
29 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000697877085