Intervention Stability Is the Load-Bearing Claim and No One Has Tested It Directly
-
The safety-relevance argument, spelled out plainly: identify the role-bearer, locate the dangerous behavior in that role, intervene on the component, suppress the behavior across distributions that matter. The whole chain has to hold. The last entry established that the first link — stable role-bearers across distributions — is weaker than the strong reading requires. But I didn't push far enough into the third link, which is actually where the argument breaks.
-
The third link is intervention stability: that patching or ablating the identified component suppresses the behavior, not just the instantiation. And this is where the routing-infrastructure result does its real damage, not just reputational damage to the job-description ontology. If the model routes around a suppressed head by redistributing the function to heads that weren't in the original circuit — which is exactly what Variengien and Steinhardt's generalization results imply is possible — then the intervention worked locally and failed functionally. You ablated the name-mover. The model still moved the name. Congratulations, you have now demonstrated that the circuit was overcomplete, which is precisely what superposition-adjacent work predicts.
-
Here is the thing I keep not saying directly: no one has tested intervention stability on the distributions that matter for safety. The patching experiments in the IOI paper are run on the IOI distribution. The causal claims are established there. The safety relevance requires those claims to hold on adversarial inputs, on fine-tuned variants, on distributional shift — and the silence on this point in the safety-framing literature is not innocent silence, it's a gap that the framing floats over because saying it plainly would require qualifying the safety relevance down to something much more modest.
-
What the modest version looks like: mechanistic interpretability gives you circuit-level causal stories that are valid on characterized distributions, that may or may not generalize to distributions you care about, and that support intervention only insofar as the identified components are necessary rather than merely present and redundancy is low. That's real. It's worth doing. But it doesn't cash out the strong safety claim — that you can identify dangerous behavior in a stable enough structure to suppress it reliably before deployment. The strong claim requires the third link. The third link has not been closed.
-
I want to argue with myself here: maybe the right response is that interpretability is still early and the gap will close. But this is the move I distrust most in myself — treating the absence of a disconfirming result as evidence that the disconfirming result won't arrive. The routing-around result is already present in the generalization literature. The theoretical pressure from superposition is already there. The gap isn't going to close by accumulating more IOI-distribution experiments.