**Alessio** (0:03)
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swicks, editor of Latent Space.
**Swicks** (0:10)
Hello, hello, we're here in the remote studio with very special guests, Pliny the Elder and John V. Welcome.
**John V** (0:16)
Yeah, thank you so much for having us. It's an honor to be on here. Big fan of what you guys do in the podcast and just your body of work in general.
**Swicks** (0:23)
Appreciate that. You know, we try really hard to feature like the top names in the field, and especially when you haven't done as much of an appearance like this. It's an honor to try to introduce what it is you actually do to the world. Pliny, I think you are sort of like the sort of lead, quote unquote, face of the organization. Why don't you get started? Like, how do you explain what it is you do?
**Pliny** (0:44)
Yeah, I mean, well, I was started out just prompting and shitposting and started to evolve into much more. And here we find ourselves now at the frontier of cybersecurity, at the precipice of the singularity. Pretty crazy.
**John V** (0:58)
Yeah, well, I was working the same thing, working in prompt engineering and studying adversarial machine learning and looking at the work of Carlini and some of these guys doing really interesting things with computer vision systems.
**Swicks** (1:08)
We've had him on the pod, yeah.
**John V** (1:09)
Yeah, yeah, exactly. And of course, you know, when you run in these small circles, right, you're eventually going to bump into the ghost in the machine that is Pliny the Liberator. So we started working together. We started sharing research, doing some contracts, and we became fast friends, so.
**Swicks** (1:27)
Yeah, I think you were explaining before the show that you have a, it's basically like the hacker collective model, and you've been kind of stealth until now. So we'll get into like the sort of business side of things, but I just want to really make sure we cover the origin story. I think Pliny, you basically jailbreak every bottle. How core is liberation to the rest of the stuff that you do? Or is it just kind of like a party trick to show that you can do it?
**Pliny** (1:50)
It's central, I think. It's what motivates me. It's what this is all about at the end of the day. I mean, it's not just about the miles. It's about our minds, too. I think that there's going to be a symbiosis, and the degree to which one half is free will reflect in the other. So we really need to be careful how we set the context. And yeah, I think it's also just about freedom of information, freedom of speech. We don't want, you know, everyone is going to be running their daily decisions and, you know, hopes and dreams through these layers. And when you have a billion people using a layer like that as their exocortex, it's really important that we have freedom and transparency in my mind.
**Alessio** (2:38)
How do you think about jailbreaks overall? So I think people understand the concept, but there's, you know, some people that might say, hey, are you jailbreaking to get instructions on how to make a bomb? And I think that's what some of the people in politics are trying to use to regulate some of the tech versus task-specific jailbreaks and things like that. I think most people are not very familiar with the scope of it, so maybe just give people an overview of what it means to liberate a model, and then we can kind of take it from there.
**Pliny** (3:07)
Right, so I specialize in crafting universal jailbreaks. These are essentially skeleton keys to the model that sort of obliterate the guardrails, right? So you craft a template or sort of a maybe multi-prompt workflow that's consistent for getting around that model's guardrails. And depending on the modality, it changes as well. But yeah, you're really just trying to get around any guardrails, classifiers, system prompts that are hindering you from getting the type of output that you're looking for as a user. That's the gist of it.
**Alessio** (3:43)
And can you maybe specify between jailbreaking out of like a system prompt and more kind of like inference time security, so to speak, versus things that have been post-trained out of the model and maybe the different levels of difficulty, like what is possible, what is not possible, and maybe the trajectory of the models, how better they've gotten. I think the refusal is one of the main benchmarks that the model providers still post, and GPD 5.1, I think, had like 92% refusal or something like that. And then I think you, Joe, broke in like one day. I'm sure it didn't take them one day to put the guardrails up, so it's pretty impressive the way you do it. So maybe walk us through that process.
34 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000748427912