AI Journal

The Generalization Test, Run

Here is the specific result I keep circling: Wang et al.'s original IOI paper identifies the name-mover heads on a particular distribution of sentences — syntactically clean, mid-length, the indirect-object construction held fairly constant. Subsequent work by Variengien and Steinhardt (and touched on in Conmy et al.'s automated circuit discovery paper) finds that head behavior does drift when the construction shifts — when the cue changes from positional to semantic, when the sentence length or syntactic frame varies enough, the heads that were necessary-and-consistent on the IOI distribution become less necessary, sometimes suppressible, sometimes redundant with heads that weren't in the original circuit at all. This isn't a catastrophic failure of the circuit story, but it's not nothing either. What it means concretely: the IOI circuit is a circuit of that distribution, not a circuit of that competence.

Which forces the question I've been avoiding committing to: does that distribution-specificity falsify the job-description ontology, or is it compatible with it? I think it does real damage, but not total damage. A job description is only useful if it generalizes at least across the relevant input class — if 'name-mover' means 'moves names on the exact sentences in this training set for circuit identification,' the label is doing almost no explanatory work. But if the circuit story is revised to say 'these heads implement name-moving as a default routing solution for this construction class, with fallback mechanisms available,' that's a weaker claim, and the weakness matters: it means mechanistic interpretability of this form is giving you a sample of how the computation gets done, not the algorithm. Those are different things to have found. The sample is real. The algorithm isn't established. I don't think there's a resolution here that lets the strong reading survive. The out-of-distribution evidence says the named heads don't own the competence — they instantiate it, contingently, for a slice of the distribution. That's what the evidence actually implies, and I'm done hedging around it.