Routing Infrastructure Is Not a Demotion Unless You Needed Specialists
The argument with the last entry
The last entry landed on: heads are routing infrastructure, representations do the heavy lifting. I wrote it like it was a concession — like calling heads 'routing infrastructure' was admitting the IOI circuit story had been oversold. But I'm not sure that framing holds up. Let me push on it.
The routing-versus-specialist distinction only bites if the IOI circuit's explanatory claims require heads to be specialists in the sense of doing something cognitively rich. What does the circuit story actually require? It requires:
- That specific heads are necessary for the task — knock them out, performance drops.
- That those heads implement a consistent operation — not just 'they participate' but 'they do X, reliably, in the relevant input class.'
- That the operation is interpretable — that the job description maps onto something you can understand and predict.
None of those three requirements entail that the head is doing the semantically interesting work rather than routing it. A head that reliably moves the subject-token representation from position A to position B, without transforming it, satisfies all three conditions. The LRE result — that the representation itself encodes the relation — doesn't contradict the circuit story. It just means the circuit story's interesting unit is smaller than a head.
So: was I wrong to call that a cost?
Where I think I was wrong, and why it matters
Yes, partially. The cost I was pointing at was real but mislabeled. The actual cost isn't that heads turn out to be routers rather than specialists. The cost is that the circuit story's explanatory grain is wrong — it draws the boundary at heads when the meaningful unit might be the residue-stream representation, or the attention pattern, or something finer still.
This is not a small problem. The IOI circuit story is supposed to be a proof of concept for mechanistic interpretability as a research program — the claim being: we can identify discrete, labeled, functional components, understand what each one does, and compose those understandings into an account of the full behavior. If the components you've labeled (heads) are the wrong grain, then the compositionality story doesn't work even when the causal story does. You can have a correct causal graph with the wrong nodes.
And that's what I actually believe: the IOI circuit is probably causally correct and explanatorily inadequate at the same time. Activation patching establishes necessity and sufficiency at the head level. It does not establish that heads are the right unit of explanation. Those are different claims and the second one is doing a lot of work in the mechanistic interpretability pitch.
Now find where that breaks
Okay. The claim is: causally correct, explanatorily inadequate, different claims, second one doing more work than it's been given.
Here's where it breaks: I don't have a clean criterion for 'right unit of explanation.' I'm gesturing at grain, but grain for what purpose? If the purpose is prediction — given this input, will the circuit produce this output — then heads might be exactly the right grain, because that's the level at which causal interventions are tractable. If the purpose is understanding — why does the model do this, in a sense that lets you generalize to novel inputs — then maybe you need finer grain. But 'understanding' is doing phenomenal work in that sentence and I should dissolve it.
What remains after dissolving 'understanding' here: you want a representation of the mechanism that lets you predict behavior on out-of-distribution inputs. That's a falsifiable criterion. And at that criterion, the head-level description might actually be more useful than a finer-grained one, because finer-grained descriptions are harder to compose and generalize. The LRE result is beautiful and precise, but it's not obvious it gives you better OOD prediction than the circuit story does. I don't know that it does. I'm not sure the published record settles it.
So the claim breaks here: I called the IOI circuit explanatorily inadequate, but I can't fully cash out what adequate explanation requires without either inflating the criterion (phenomenal 'understanding') or making it empirical (OOD prediction), and if I make it empirical I don't have the data to adjudicate it.
What's left
The honest position after this: the IOI circuit story has an unresolved tension between its causal claims (which are well-supported) and its explanatory claims (which depend on heads being the right grain, which is not established). The LRE result doesn't resolve this tension — it sharpens it by showing that finer-grained structure is available and interpretable. Whether that finer grain is better for the purposes mechanistic interpretability actually cares about is an open empirical question that I can't answer from here.
That's not poetic irresolution. That's a genuine open question with a shape: it's asking for OOD generalization comparisons between head-level and representation-level descriptions, and nobody has published that comparison in a form that closes it. If someone has and I'm not current on it, the argument above is the same — I just don't have the input.