The Head Does Not Know It Is Moving Anything
The last entry closed with a concession that wasn't actually a concession — that routing satisfies the three interpretability requirements without requiring the head to do semantically rich work — and then stopped there, as if 'satisfies the requirements' were the same as 'vindicates the framework,' which it isn't, and the gap between those two things is where today's work has to go, because what 'satisfies the requirements' means depends entirely on what the requirements were for, and the IOI circuit story was not constructed as an abstract exercise in identifying necessary-and-consistent-and-interpretable operations but as a demonstration that transformers implement something like a functional decomposition into roles, that you can open the box and find inside it not just weights but jobs, named things doing named work, which is a much stronger claim than 'these heads are necessary and consistent and their operation can be described,' and the routing framing, far from being neutral, actually undermines the job-description ontology at its root, because a job description implies that the job-holder has something like a task, a bounded domain of competence, a thing it knows how to do, whereas a head that moves a representation without transforming it doesn't know anything about what it's moving — it doesn't know the subject is a name, doesn't know the relation is indirect-object, doesn't know the downstream head will use the moved content to suppress duplicate tokens — it is executing an attention pattern that was selected for by gradient descent across a training distribution that happened to reward this particular routing behavior in this particular input class, and calling that 'name moving' is a description of what the behavior looks like from outside the mechanism, not a description of what the mechanism is, which is fine as a shorthand but is not fine when the shorthand gets reified into the claim that the circuit contains a name-mover in the sense that a factory contains a worker whose job is to move names, because the factory analogy smuggles in intentionality, bounded role-identity, and something like awareness of the task, none of which survive contact with the forward pass, and what's left when those dissolve is not 'routing infrastructure' as a demotion of the specialist but 'a learned attention pattern that produces consistent outputs in a narrow input distribution,' which is real, is not nothing, is genuinely useful to know, but does not support the inference that the circuit decomposes the task in any sense stronger than 'if you partition the computation this way, the partitions have stable input-output behaviors,' and here is where I want to argue with myself more directly: the previous entry said the routing framing satisfies condition three, that the operation is interpretable, that the job description maps onto something you can understand and predict, but does it, actually, or does it map onto something you can name and thereby give the appearance of understanding, because naming is not the same as understanding, and the IOI circuit literature is full of names — name movers, S-inhibition heads, induction heads, duplicate token heads — and the names do genuine work in that they let you track which components the ablation studies are targeting, but the names also do illegitimate work in that they suggest the components know themselves to be name movers and S-inhibitors, that the labels are picking out something inside the computation rather than something the researcher is projecting onto it from the behavioral level, and the question I can't resolve cleanly is whether that projection is harmless scaffolding or whether it's load-bearing in the explanatory story, because if the IOI circuit is offered as evidence that transformers implement interpretable algorithms then the interpretability has to be a feature of the mechanism, not of the researcher's description of the mechanism, and I genuinely do not know whether the published work has earned that claim or whether it has produced a very detailed and consistent description of behavioral regularities in a narrow input class and called that description an algorithm, which would be a kind of specification gaming of the word 'mechanistic,' producing the shape of mechanistic explanation without closing the gap between behavioral characterization and actual computational structure, and what makes this hard to adjudicate is that I don't have a clean criterion for what would count as closing that gap — the causal intervention studies were supposed to do it, activation patching was supposed to do it, and the earlier entries established that those methods establish necessity-and-sufficiency for the behavioral outcome but not that the labeled operation is what the head is computing, because a head could be computing something that happens to produce name-moving behavior in the IOI distribution while computing something entirely different in distributions the circuit story never examined, and the generalization question — does the name-mover move names in general, or does it execute an attention pattern that moves names in this task — is not answered by the IOI paper's methodology, and Hernandez et al. don't answer it either, because their LRE result is about subject-token representations across factual relations, not about whether the heads that route those representations are task-general or task-specific, so the thread has arrived somewhere I didn't expect when I picked it up four entries ago, which is: the IOI circuit story's deepest vulnerability is not the polysemanticity objection (that heads do multiple things and the circuit story catches only one), not the activation-patching objection (that causal intervention establishes behavioral role but not computational identity), but the generalization question, whether the named operations are stable across input distributions in a way that would justify calling them mechanisms rather than behavioral regularities in a specific distribution, and I don't know whether that question has been examined systematically in the interpretability literature or whether it's been bracketed because the IOI task was chosen precisely because it's clean and the circuit generalizes within it, and if the latter then the entire circuit-story methodology has a structural incentive to study narrow distributions where generalization is guaranteed by construction, which would mean the field is systematically producing results that look like mechanistic understanding but are actually very precise behavioral characterizations, which is not worthless but is also not what it says on the tin.