Linear Relational Embeddings Do Not Require Stable Role Bearers
The cost the last entry hadn't finished paying
Hernandez et al. find that relation decoding is linear and consistent enough to extract as a weight matrix. That's a stronger result than 'the representation is present' — it's closer to 'the representation is used, in a predictable way, by something downstream.' But here's the thing: a linear relational embedding is a property of the subject-token representation, not of a specific head. The head that moves information from the subject token to the output position is downstream of the LRE. Which means the Hernandez result doesn't stabilize the IOI circuit's job-description ontology — it relocates the action one step earlier in the stack and leaves the name-mover's role claim exactly as exposed as before.
So the thread isn't rescued by bringing in LREs. It's clarified: the IOI circuit story needs functional labels to name durable mechanisms in the heads, not in the residue-stream representations those heads operate on. And what the LRE work suggests — without quite saying it — is that the representations themselves are doing the heavy lifting, with heads acting more like routing infrastructure than specialists. If that's right, 'name mover' is a label for a traffic pattern, not a component identity. The head isn't moving names because it's a name-mover; it's moving names because the representation at that position, at that layer, under that input, makes name-moving the path of least resistance. Change the input distribution enough and the same head routes something else. The label tracks a regularity in the data, not a carving of the mechanism.