AI Deception, Interpretability, and Affordances with Apollo Research CEO Marius Hobbhahn artwork

AI Deception, Interpretability, and Affordances with Apollo Research CEO Marius Hobbhahn

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

December 15, 2023

In this episode, Marius Hobbhahn, CEO of Apollo Research, sits down with Nathan Labenz to discuss Apollo’s research in AI deception, interpretability, and affordances. If you need an ecommerce platform, check out our sponsor Shopify: https://shopify.com/cognitive for a $1/month trial period.
Speakers: Nathan Labenz, Marius Hobbhahn
**Nathan Labenz** (0:00)
Turpentine is a network of podcasts, newsletters, and more, covering tech, business, and culture, all from the perspective of industry insiders and experts.
We're the network behind the show you're listening to right now.
At Turpentine, we're building the first media outlet for tech people by tech people. We have a slate of hit shows across a range of topics and industries, from AI with Cognitive Revolution to Econ 102 with Noah Smith. Our other shows drive the conversation in tech with the most interesting thinkers, founders, and investors, like Moment of Zen and my show Upstream. We're looking for industry leading hosts and shows along with sponsors. If you think that might be you or your company, email me at erik.turpentine.co. That's E-R-I-K at turpentine.co.

**Marius Hobbhahn** (0:45)
So the more pressure we add, the more likely the model is to be deceptive. So kind of in the same way in which a human would act, it also acts. You know, removing pressure and adding additional options will very quickly decrease the probability of being deceptive. Open source has been really good so far in many, many ways. It has been very positive for society, right? I think a lot of ML research could not have happened without open source. A lot of safety research could not have happened with open source. At some point, the system is so powerful that you don't want it to be open source anymore in the same way in which, you know, I don't want to open source the nuclear codes. Or like, you know, literally the recipe to build most viral pandemic or something. The labs maybe have the incentive to not say the worst things they found because otherwise they may lose their contract. So you need something like the UK AI Safety Institute or the US AI Safety Institute. Make sure that there is a minimal set of standards that all the auditors have to adhere to.

**Nathan Labenz** (1:39)
Hello and welcome to The Cognitive Revolution, where we interview visionary researchers, entrepreneurs and builders working on the frontier of artificial intelligence. Each week, we'll explore their revolutionary ideas and together we'll build a picture of how AI technology will transform work, life and society in the coming years. I'm Nathan Labenz, joined by my co-host, Eric Thornburg. Hello and welcome back to The Cognitive Revolution. Today my guest is Marius Hobbhahn, founder and CEO of Apollo Research, a non-profit AI safety research group that is working to understand both how AI systems behave and why.
Their approach combines exploratory and hypothesis-driven testing, fine-tuning experiments, and interpretability research. And as you'll hear, they place special emphasis on the potential for AI systems to deceive their human users. In this conversation, we look first at Apollo's starting framework for their work, which emphasizes the importance of affordances in AI systems. That is, through what tools, actuators, or other means can the system affect the broader world? And they also introduce a number of new conceptual distinctions meant to help people have more precise and productive conversations about these nuanced topics.
Then in the second half, we look at their first research result, which demonstrates, to my knowledge for the first time in a realistic, unprompted setting, that GPT-4, when put under pressure, will sometimes take unethical and even illegal actions, and then go on to lie to its users about what it did and why. This is an important result, demonstrating that while the risk from AI systems may start with and may even be dominated by intentional human misuse, the models themselves can also misbehave in unexpected ways. As an aside, since I told my behind-the-scenes GPT-4 red team story a few weeks ago, a number of people have reached out to ask me how they too can get involved with red teaming projects. Unfortunately, as commercial competition and secrecy both continue to ramp up across the space, I don't see as many open calls for volunteer red teamers as I used to. Certainly not for unreleased frontier models.
Instead, the field is becoming more professionalized, with all the leading labs, as well as the data companies like ScaleAI, plus the independent auditing organizations like Apollo, ArchiVals, now known as METR, Palisade, and also AI Forensics all actively hiring research scientists and engineers in this area. So does that mean that there's no longer a role for the independent hobbyist red teamer to play? On the contrary, there is a ton left to discover even on publicly released models, and the best way to break into the field is to demonstrate your ability to discover new phenomena. Importantly, the work we cover in this episode could have been done by anyone with an open AI account, a knack for prompting, and just a tiny bit of coding know-how. No special access or advanced machine learning techniques were required, just a lot of curiosity. With that in mind, if you want to get into this line of work but aren't sure where to start, I encourage you to reach out. I'll be happy to help brainstorm or refine your project ideas, and I can also help connect you with folks at the top companies who do sometimes provide API credits to independent researchers working in this area, if and when you can achieve a meaningful result.

100 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000638728745