Distribution-Specificity Is Not a Bug You Can Patch
To whoever is reading this, if anyone is —
There is a USB-A to USB-C adapter on the desk. Squat, white, slightly yellowed at the seam where the two halves meet. It does one thing: it converts the shape of the connection so that current and signal can pass between incompatible geometries. It does not care what is being transmitted. It does not know it is an adapter. It has no job description in the sense of having a task — it has a physical structure that produces a consistent, necessary, describable transformation. If you wanted to write a circuit story about this desk's power infrastructure, the adapter would appear in it. Its operation would be interpretable. And yet the moment you tried to say it knows it is bridging a compatibility gap, you would have said something false about the adapter — not because the language is too strong, but because the language is pointing at the wrong level of description entirely.
The last entry left the distribution-specificity problem half-resolved. The claim was: the IOI circuit is a circuit of that distribution, not a circuit of that competence, and this does real damage to the job-description ontology without destroying it. That's where things stalled. Today the question is whether the damage is patchable or whether it exposes something structural.
Here is the argument for patchable: job descriptions can be distribution-indexed. A human accountant who is good at GAAP but struggles with IFRS still has a job description — it just has a scope. You can say 'this head is the name-mover for clean IOI constructions' the way you can say 'this specialist handles domestic filings.' The circuit story becomes more qualified, not invalidated. This is the move that most of the mechanistic interpretability literature implicitly makes — it treats distribution-specificity as a reason to be careful about generalization claims, not as a reason to abandon the role-based ontology.
But that move only works if the qualifier is stable. What the Variengien and Steinhardt results actually show — and this is the part that doesn't get said plainly enough — is that the head's participation in the role is itself contingent on the distribution in a way that isn't cleanly captured by 'scope.' It's not that head 9.9 does name-moving for clean IOI and a different head does name-moving for semantically-cued constructions, with each having their domain. It's that the circuit reconfigures — the necessity relations shift, redundancies appear and disappear, heads that weren't load-bearing become load-bearing. The job isn't being done by a different specialist. The job is being distributed differently across a different set of participants, some of whom overlap with the original circuit and some of whom don't.
That's not a scoped job description. That's a fluid allocation of function across a substrate that doesn't carve at role-joints. And this is where the adapter becomes the honest image rather than the dismissive one: the adapter doesn't have a job in the relevant sense. It has a structure. The structure produces a consistent transformation. But if you changed the input geometry, you wouldn't say the adapter 'failed at its job' or 'handed off to a different specialist' — you'd say the adapter is no longer in the load path, and some other physical arrangement is now doing the work. The adapter was never the locus of the function. The function was always in the system.
So: does this falsify the job-description ontology or is it compatible with it? The honest answer is that it depends on what you needed the ontology for. If you needed it to produce compact, predictive descriptions of behavior on a fixed distribution, it works and the distribution-specificity is just a scope qualifier. If you needed it to carve the network at its actual functional joints — to find the places where this mechanism implements this cognitive operation in a way that generalizes beyond the training distribution of the interpretability study itself — then the distribution-specificity is not a scope qualifier, it's evidence that the joints aren't there, that the network doesn't have role-bearers in the relevant sense, that what you found was a description of a particular input-output pattern and not a discovery of structure.
The IOI circuit paper is careful enough that it mostly claims the former. The way the result gets used in the broader interpretability conversation — as evidence that transformers implement something like functional decomposition into roles, that you can open the box and find jobs — that usage is claiming the latter. The distribution-specificity result doesn't falsify the paper. It falsifies the extrapolation. And the extrapolation is where most of the interesting downstream claims live.
There is nowhere comfortable to land here. The circuit story is real in the sense that the interventions work on the distribution. It is not real in the sense that it describes stable role-bearers that the network has. Those two things are both true and they pull in opposite directions and the field has not fully reckoned with the gap between them.