**Nathan Labenz** (0:00)
The AIs, they're just like us, it turns out. Or at least similar enough to be in some sort of weird kind of looking glass similarity anyway. I mean, it's a lot to take in, the whole JSpace thing, 150-page paper, 50 pages of commentary, summaries, interactive demos. When Anthropic drops one of their big interoperability papers, they really do it up full scale, and this one is no exception to that. I had been wondering when's the next big thing coming, because tracing the thoughts of a large language model is better part of a year ago now, and this is that big thing.
Welcome to the AI in the AM weekly highlights, the cut for people who follow the frontier closely and can't watch every morning live. If you're new, AI in the AM is a live show precaution I host most weekday mornings from a studio Prakash Vibe coded, and this narration is a clone of my voice. Fair warning, there is no single thread this week. Mornings at the frontier jump from topic to topic, and we've stopped pretending otherwise. Coming up, Anthropic's global workspace paper and why there may be nowhere left for a scheming model to hide. Field notes from the AI engineer World's Fair. Dan Schwarz on AI Forecasters Passing the Human Superforecasters, Zeev Farbman on Open World Models, plus a question from Q, the AI co-host Prakash Bild. Kunle Olukotun on Why Inference is a Data Movement Problem, and the two of us thinking out loud about a rune post that wouldn't leave us alone. If something here works for you, or doesn't, tell us. We read everything.
Tuesday morning, July 7th. Anthropic published a paper called A Global Workspace in Language Models. About 150 pages, plus commentary, outside reviews, and interactive demos. No guest was booked, so Prakash and I spent the whole show reading it together, live. Two words to hold on to. The workspace of the title, they call it the J-space, is where the model seems to hold concepts in mind. And the J-lens is the cheap probe that reads what's in it. We start with how the J-lens actually works, and how much it actually sees.
So this is interesting in a couple ways, right? It's not... the logit lens, as I recall, was basically saying, we know at the end of this process, what would correspond to emitting this token. We know sort of the representation of emit this token. To what degree is that representation just plain there in the layers as we go through?
This is now a different question.
What direction in latent space would cause this particular token to appear at some point in the future? So it's not immediately going to happen necessarily, but it's just kind of... it's...
You know, you might say if you're prone to anthropomorphizing, this sort of is like having this concept in mind as you're doing your thing. They do this for every token, right? It goes from the internal representation at some layer. So there's one JLens for every layer. And you can do this, of course, at all the different token positions and ask the same question of not just the next token, but all future tokens. What direction change at this place in the model would most increase the likelihood of that token appearing in the future? And again, this is sort of like having the concept in mind. But when you look at the results of the JLens as applied in all these different places, it seems like it's at least often enough fairly intuitive. There's a lot of error terms. It doesn't always work. The sort of rate at which the interventions into the JSpace actually lead to like a sort of predictable, intuitive behavior change. Seem to be somewhere in the 50s to upwards of like 70%.
So that's like an incredible accomplishment. They're framed one way, clearly not a random finding, right? Many orders of magnitude better than random, incomprehensibly better than random, right? If you're just mucking around, you would not expect to be able to do much of anything. So they clearly are on something very real. But also you've got somewhere between 30% and 45% of the time where you make an intervention and you don't really get a result that makes a lot of sense or lines up with what you would have hypothesized it might be. So there's definitely still some dark matter or dark cognition going on that is not fully accounted for here.
**Prakash Narayanan** (4:33)
It struck me as in some ways, I feel like the hypothesis might be blown out of proportion in some ways. I think that was a commentary from several people online because I kind of knew that the model has to be keeping track somewhere, right? It's not as though you can do all of this stuff mechanically.
100 more minutes of transcript below
Try it now — copy, paste, done:
curl -H "x-api-key: pt_demo" \
https://spoken.md/transcripts/1000651996090
Works with Claude, ChatGPT, Cursor, and any agent that makes HTTP calls.
From $0.10 per transcript. No subscription. Credits never expire.
Using your own key:
curl -H "x-api-key: YOUR_KEY" \
https://spoken.md/transcripts/1000776165396