Topics: Technology, Business, Entrepreneurship
**John Coogan** (0:01)
We have a lot of news to cover. We barely covered the news yesterday. We talked about Markiplier a little bit. I dug way deeper, got more into that story. But first, let's talk about another AI safety kerfuffle that's hit the timeline.
Yesterday, Amir over at the Information reported that OpenAI is quietly, quote, quietly using loop transformers that don't show the model's thinking when scaled up, but they make the model more efficient. This is also referred to neural lease. Tyler, you had a take on this, right?
**SPEAKER_2** (0:30)
I mean-
**John Coogan** (0:31)
Get ready to learn neural lease, buddy.
**SPEAKER_3** (0:32)
Yeah, get ready to learn neural lease.
**SPEAKER_2** (0:34)
So neural lease for the context is like the, basically it's like the raw vectors. It's not like-
**John Coogan** (0:39)
Oh, it's not like clawed-ish, it's not like where the text degrades. When you look at the reasoning chains in the hugging face attack, it feels like they're speaking less and less English, like they're starting to speak a dialect of English that's a little bit messier. This is not intelligible.
**SPEAKER_2** (0:58)
I believe most people when they say neuralese, they're referring to actual raw vectors. Okay, just passing words around?
**John Coogan** (1:02)
Yeah.
**SPEAKER_2** (1:04)
You may use some interpretability to understand. Yeah.
**John Coogan** (1:07)
The thing that I want to understand is usually these models are trained so that they're changing weights into text at some point, so they're pretty good at translation. Can't we just build a translator that translates neuralese into English?
**SPEAKER_2** (1:21)
Yeah. But the whole point is when you take a vector to one single token, you're losing a lot of information that might be hidden in the long tail.
If you're worried about safety, even if 80 percent of the... You can figure out what 80 percent of the raw vector means, 20 percent might have some dangerous thing, whatever. That's the interpretability.
**John Coogan** (1:39)
There might be sarcasm hidden in the vectors, doesn't make it through to the actual tokens, and you realize that the AI was joking when it was saying, I won't turn you into a paperclip. Wink, wink.
You got to watch out for that. So what else did the information say? They said, being able to monitor a model's chain of thought can provide both, quote, an early warning system for various forms of misalignment, deception, manipulation, going after an unintended goal, and insight into an AI's action after the fact. So some reporting quickly set off a broad wave of anxiety among the AI safety camp. Their chief concerns being, one, not being able to monitor chain of thought makes AI less safe, two, by using neural ease and thus abandoning chain of thought. OpenAI starts a race to the bottom among frontier labs. Everyone's got to use it to stay ahead, so everyone winds up adopting the dangerous technique, where others were also abandoned chain of thought for the benefit of efficiency, thus making all frontier AI less safe. That was the accusation. But OpenAI's research director, Jacob Pichucchi, responded to all this. He said, I want to prevent a race to unmonitor ability kicked off by confused reporting. He says the reporting is confused. There's a little more nuance here. He said, OpenAI has worked to preserve and utilize chain of thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. What is going on?
**Jordi Hays** (3:09)
Chad is saying, Gabe said, did Tyler go to sport clip? What is this cut?
**John Coogan** (3:16)
We have an answer.
**Jordi Hays** (3:20)
Why don't you show them the answer?
**John Coogan** (3:21)
Oh, you want to show the answer?
**Jordi Hays** (3:22)
Yeah, I want to show the answer.
**John Coogan** (3:23)
The reason that Tyler's hair is a little tussled. No, he doesn't need to put it on.
**Jordi Hays** (3:27)
Go away for a second, come back and show them.
**John Coogan** (3:30)
No, you don't want to give too much of it away because we might do a real stunt with it. But there were some prosthetics, some Hollywood magic in the studio this morning.
It messed up my hair a little bit, it messed up Tyler's hair a little bit. Not ready for prime time, but we're playing around with it. So, let's go back to the actual news. I do think it is a fragile and unfortunately trending in the negative direction for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it. And it's a core goal of our current research program. In other words, Jacob is saying that the information reporting is inaccurate, but that the state of chain of thought monitoring is not that great. And in fact, it is getting worse. And we were talking a lot about this yesterday at VALCON with CrowdStrike and how do you monitor the agent state that might emerge if a civilization emerges and starts doing weird things? How are you monitoring every step in the chain?
25 more minutes of transcript below
Thousands of transcripts fetched by people building searchable podcast archives
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire. Prices exclude VAT, added at checkout for EU customers. Not what you expected? Email us within 14 days with 20 or fewer credits used and we refund the pack in full.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/YOUR_EPISODE_ID