Topics: Tech News, News, Technology
**Connor Leahy** (0:00)
Fundamentally, technology is power. Power isn't good or evil, it's power. It's amplifying. If you have uncontrolled power, you have chaos. If you add power to a chaotic system, you don't get a nice, orderly, nice world. You get a more chaotic, more dangerous, more violent, more unjust world.
**Laurie Segall** (0:21)
I'm Laurie Segall, and you're listening to Mostly Human, a tech podcast through a human lens. Connor, welcome.
**Connor Leahy** (0:29)
Thanks so much for having me.
**Laurie Segall** (0:31)
Let's start. I want to get into the incidents and what happened at OpenAI, Anthropic, with these models going rogue, but I would love to just talk a little bit before we get into it about your background.
**Connor Leahy** (0:41)
So I'm Connor. I'm originally from California, but I grew up a lot in Germany. My mom's German and in a small German town back in the day when I was a teenager, I was thinking about how to make the world a better place. How could I solve many problems? And so I thought, well, if I could just solve intelligence and AI, well, we could solve all these problems, you know? We could cure all these diseases. We could maybe we could solve climate change. It'd be great. How hard could it be? What's the worst that could happen? And so as I got a little bit older, I started realizing, well, if you build something that was so smart that it could do something like cure all diseases, you know, something that our greatest scientists for many generations have been working on, that's a very powerful thing and potentially a very dangerous thing. How do you even control something like that? What would it mean to control something so powerful? So this became the problem that I dedicated my career to solving. So I have a technical background.
I worked a lot at various startups, but I also founded an open source group where we built some of the first open source large language models in the world.
There's a group called Luther AI, so I did a lot of open source work, did a lot of technical work. I ran a company where we worked on AI safety and control, did a lot of RID in that direction. And recently, I have moved on from technical work to join as the US Executive Director of ControlAI, which is a non-profit advocacy organization, which is pushing for the prohibition on the development, which I'm sure we can talk about what is called super intelligence. So advanced AI systems that we cannot control.
**Laurie Segall** (2:15)
Fascinating. And so much to unpack there. And I think it's such an interesting point if they can do all these incredible things, what are the bad things that it can do. And so I would say as someone who is really focused on safety and guardrails and control of your whole career, I was really curious to bring you in because something happened recently that I was just like, well, I feel like this was inevitable, but came a little quicker than I thought it would, and this doesn't seem good. And the thing I'm describing is OpenAI recently announced that its AI models secretly broke out of a secure testing environment and hacked into another AI company called Hugging Face, which you can get into and talk to us about what that is, in order to cheat on an evaluation. So can you walk us through what happened? I mean, the headline is AI went rogue.
So what happened and what do you make of it?
**Connor Leahy** (3:12)
Yeah, I can basically only agree with your assessment there, is that this is something that was predictable and it's still crazy to see it happen, especially this early. I think that's really important to understand before I explain exactly what happened here, is that modern AIs are not chatbots. Originally, like GPT and the original ChatGPT were what are called large language models. They were trained to predict text. You gave them all this text off the Internet and their task was to predict the next word. This is not how modern ChatGPT works. There is some component of this. It's a very important component of the system, but it is not how these systems are trained anymore. These systems nowadays are what are called agents, and they are trained with a process that is called reinforcement learning, in addition to the large language model component.
You can kind of imagine that there's two steps, so to speak. In the first step, you do this large language modeling. So you train them just to predict a bunch of text. This makes them pick up a lot of knowledge. Like they now know a lot of things about how the world works, about how sentences work, about how everything is on Wikipedia, whatever. So they pick up a bunch of useful stuff.
65 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID