Exploitable by Default: Vulnerabilities in GPT-4 APIs and “Superhuman” Go AIs with Adam Gleave of Far.ai artwork

Exploitable by Default: Vulnerabilities in GPT-4 APIs and “Superhuman” Go AIs with Adam Gleave of Far.ai

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

March 27, 2024

In this episode, Nathan sits down with Adam Gleave, founder of Far AI, for a masterclass on AI exploitability.
Speakers: Erik Torenberg, Adam Gleave, Nathan Labenz
**Erik Torenberg** (0:01)
Hey, everyone, Eric here. We're really excited about a new AI show from Turpentine called Autopilot, hosted by Will Summerlin.
This podcast explores the adoption and rollout of AI in the industries that drive the economy, and the dynamic tech founders bringing rapid scalable change to slow moving industries. From law, to hardware, to aviation, we'll interviews founders backed by Benchmark, Greylock, YC, and more to learn how they're automating at the frontiers and entrenched industries. Click on the link in the description to subscribe to Autopilot.

**Adam Gleave** (0:31)
This photo experiment of like what could Einstein's brain and a vats do. It's like, look, you can't take over the world and no matter how smart you are, if all you can do is just think. And now, well, we're not just letting models think, we're giving them access to run code, to spin up virtual machines, to execute external APIs. So I think this should be sort of a big part of your threat model and part of your evaluation for the safety of the system is not just how capable is it, but also like what does it actually have access to and increasingly we're giving it access to more and more things. So the fact that we can maybe just about make it really hard for an attacker in Go is not much consolation when we think about actually securing Frontier General Purpose AI systems.
Zuckerberg just announced $7 billion in compute investment for Frontier models. So if they were to spend 1% of that, right, $70 million on AI safety, then that would be not far from doubling the amount of revenue being spent on AI safety.

**Nathan Labenz** (1:33)
Hello and welcome back to The Cognitive Revolution. Today, I'm excited to share my conversation with Adam Gleave, founder of Far AI. Adam and his colleagues are doing critically important work exploring the robustness and alignment of AI systems. And their results show clearly that today's machine learning systems are exploitable by default.
We begin with a discussion of Adam's recent blog post detailing a number of exploits in Gpt-4's fine-tuning and assistance APIs. Amazingly, on seemingly every dimension, there are substantial vulnerabilities. For starters, they report accidental jailbreaking via fine-tuning, a phenomena wherein a naive developer with no ill intent fine-tuning on purely benign examples often still ends up removing safety filters to their own surprise.
Purposeful fine-tuning attacks, as you'll hear, can do much more still, including generating targeted political misinformation, malicious code, and even personal information such as private email addresses. Meanwhile, the assistance API can be hijacked by malicious users and effectively turn on its host application by divulging private information from its knowledge base and even executing arbitrary function calls.
While OpenAI and others are certainly working hard on safety, the ease with which these exploits are found reflects the fact that controlling such powerful models is a fundamentally hard problem. Gaining robustness comes with a significant tax. More compute and more development time are required, and still performance is often somewhat degraded in the end. It's a hard trade-off that can't be ignored, and a real wake-up call for anyone who thinks AI systems will be safe by default.
In the second half of the conversation, we turn to Adam and team's work on superhuman go-playing AIs. In a gray box setting, which means that they could query the AIs, but not see their internal weights or states, the Far AI team was able to find strategies that reliably beat these quote-unquote superhuman systems. And in a result reminiscent of the Universal Jailbreak paper that we covered in a previous episode, they found that the unusual strategies they discovered, which advanced human players would easily defeat, did sometimes transfer to defeat other advanced go-playing AIs as well. It's a striking reminder that even ostensibly superhuman systems often have deep-seated, exploitable flaws. Looking forward, Adam is working to develop empirical scaling laws for adversarial robustness. With model capabilities improving much faster than robustness today, this kind of security mindset research is critical, because in all likelihood, closing the capabilities robustness gap will require both conceptual breakthroughs and a lot of diligent work. As always, if you find this work valuable, please share this episode with others who might appreciate it. I'd suggest this one for anyone who thinks that AI safety and control will somehow just take care of itself. And always feel free to reach out with feedback or suggestions via our website, cognitiverevolution.ai, or by messaging me on your favorite social network. Now, please enjoy this overview of the current state of AI robustness, safety, and control with Adam Gleave of Far AI. Adam Gleave, founder of Far AI, welcome to The Cognitive Revolution.

**Adam Gleave** (4:55)

95 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/1000650660638