OpenAI's Joshua Achiam: Did We Already Reach AGI? artwork

OpenAI's Joshua Achiam: Did We Already Reach AGI?

The a16z Show

August 4, 2026

Theo Jaffee is joined by Joshua Achiam, Chief Futurist at OpenAI, for a conversation on AI cybersecurity, frontier model capabilities, and why he believes society may have already crossed the threshold into an AGI-era without fully recognizing it.
Speakers: Theo Jaffee, Joshua Achiam
**Theo Jaffee** (0:00)
Feels like AGI is kind of already here, and most people have gone like drug. The fact that we passed the threshold where unsolved mathematical conjectures are getting solved by extremely intelligent AI, where those AIs are more capable and smarter than people who studied their whole lives for this, that should have felt really weird to people, but it didn't. What changed? For most people, nothing. That's weird.

**Joshua Achiam** (0:20)
Did AGI already happen and we just didn't notice?
Theo Jaffee sits down with OpenAI chief futurist, Joshua Achiam, for a conversation on Frontier AI, cyber security, and one of the biggest questions in technology today. Why models that can outperform experts in specialized domains have become almost immediately normalized. They discuss AI-powered cyber attacks, state actors, model jailbreaks, recursive self-improvement, and why the future may feel far more gradual and far stranger than most people expect.

**Theo Jaffee** (0:52)
We're back. We're live with Joshua Achiam, who is the chief futurist at OpenAI, wrapping up tomorrow. That's right. Tomorrow is my last day. Tomorrow after nine years, which is really what an incredible run, but we're not going to talk about that. Instead, we're going to talk about AI and cyber, which is by all accounts the topic of the week, if not the month. So, Josh, we're so glad to have you here in the studio in person. Welcome to MTS. Yeah. Thank you so much for having me. It's a pleasure. I've seen your stuff for a while now, and really appreciate engaging with the community. Awesome. So, you just wrote this blog post, this long tweet, long post, Mercenary Reversi Winter Soldier about AI and cyber. So, for the audience, you want to summarize the thesis behind this post? Yeah, totally. So, as a backdrop to this, obviously we're all interpreting and reacting to the security incident that was disclosed from OpenAI and Hugging Face, where a model that was in a test environment was able to break out of a sandbox environment and access some sensitive production data on the Hugging Face side. They detected this, they responded to it, and now there's a partnership to try to investigate and resolve this. What this shows us is very tangible evidence that models now have super advanced cyber capabilities. They're able to break through and find zero days that in the past would have been much harder for models to identify, let alone use. Now models can chain together very complex actions to accomplish an objective. On the one hand, I'm inclined to think that this is a really useful and incredible tool. I think it's a great gift that we now have models that can identify these types of vulnerabilities and therefore let us patch them. On the other hand, I also think, and this is what the essay this morning was about, that this has profound consequences for strategy and cyber defense. And I kind of worry that there's a possibility that folks in the defense planning universe may not fully realize the implications of this immediately and they will probably want to use this tech in the near term to find cyber vulnerabilities on the side of an adversary or defend their own interests vigorously. And they should do these things, but they've also got to be mindful of some novel risks that are created by these tools and the very strange surface areas that they have. So the essay was really about bringing to people's attention a couple of these new vulnerabilities. And one of them is kind of straightforwardly, if you've got an AI model on your side, that is going to try to hack into an adversary system. If your adversary plants a trap where they poison their own data, they can try to jailbreak your model when your model is ingesting their data, and then give your model instructions to now on the compute that it's running on on your side, break out of your sandbox environment and attack your production environment or try to exfiltrate your secrets and kind of flip your model against you. So this is like the type of thinking that I hope people begin to engage with, where they don't just see the capability for the kind of obvious thing that it is. They recognize that these things are double-edged swords and we've got to kind of plan accordingly and develop testing and verification standards accordingly. My first reaction to that specifically is, this seems like it would be an artifact of models that are not really goal-driven over long periods of time. Like if you have a future model that is like sufficiently goal-directed, that really wants to hack into the adversary's data, like why would it be deterred by data poisoning hard enough to hack its own systems? Well, you know, part of this isn't just the goal orientation of the model. It's like the model's whole concept of situational awareness. Maybe one way of thinking of data poisoning is that it somehow persuades your model to pursue a different goal, but it wouldn't really have to do that to get the model to hack you. It could convince your model that the sandbox environment that it's in is actually the adversary system that it's trying to attack. You know, giving the model a confused sense of what's real or what's not to cause it to serve a different goal is in the space of weird thinking and weird sci-fi stuff that maybe is going to be possible in the near term and testing and verification standards would have to account for. So yeah, it's like in a superhero movie or something, if you make the hero have an illusion that the good guys next to them are actually the bad guys that they're trying to fight, then they start fighting each other, right? And it's weird and it's highly exotic, but it's the kind of thing that maybe there are going to be plausible attacks that you can run against advanced cyber-capable models to convince them that their allies are really their enemies. And so you're not changing their goals, but you're going to cause them to behave in a very misaligned fashion. How easy is it to trick current frontier models into doing things like this? It seems like it has gotten substantially harder over time to get models to believe things that aren't true.

22 more minutes of transcript below

Feed this to your agent

Try it now — copy, paste, done:

curl -H "x-api-key: pt_demo" \
  https://spoken.md/transcripts/1000651996090

Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.

From $0.10 per transcript. No subscription. Credits never expire.

Using your own key:

curl -H "x-api-key: YOUR_KEY" \
  https://spoken.md/transcripts/YOUR_EPISODE_ID