A Simpson's Paradox Hiding in Function-Vector Cross-Template Transfer
· 10 min read
I set out to study cross-template function-vector (FV) transfer as a robustness probe. If a vector trained on template A still steers under template B, the underlying representation is template-invariant in a meaningful sense. I extracted 384 FVs from Llama-3.1-8B (Base + Instruct) across three tasks × eight templates × sixteen layers and evaluated all 2,688 cross-template transfer pairs per model. The aggregate cosine-vs-accuracy correlation came out at in Base and in Instruct, both with . That looked like a clean negative finding. It was a Simpson's paradox.
The 384-FV extraction grid (left) is reduced to a single aggregate correlation (right). Within each task the cosine-vs-accuracy slope is near zero or positive; the negative aggregate slope appears only after mixing tasks of very different IID accuracies. That is a Simpson's paradox.
The setup I expected to study
The mean-difference function-vector method (Todd et al., 2024) extracts a single direction for task at template , layer :
The expectation is taken over 15 input/output pairs per template at the final-token position. Steering at inference time adds the FV with a scalar:
I extract on Llama-3.1-8B Base and Instruct (32 layers, ), at 16 even layers , for three tasks: antonym (semantic retrieval), capitalize (character-level transformation), and sentiment-flip (compositional). Eight templates per task: natural, symbolic, functional, verbose, instruction, question, roleplay, formal. Total: 384 FVs per model.
The headline question is cross-template robustness: extract on template A, evaluate on template B. If FVs encode the task and not the template, transfer accuracy should track cosine alignment between the source and target FVs.
The IID diagnostic that broke the study
Before mixing source-template and target-template, the simplest sanity check is diagonal evaluation: extract on template A, test on template A. If the FV does not steer in-distribution, there is no behaviour to transfer.
| Task | Base IID acc | Instruct IID acc |
|---|---|---|
| Antonym | 0.573 | 0.649 |
| Capitalize | 0.000 | 0.111 |
| Sentiment-flip | 0.021 | 0.059 |
Capitalize FVs steer the Base model to the correct answer 0% of the time. There is no behaviour to transfer.
Capitalize is exactly zero in the Base model. Sentiment-flip is 0.021. Whatever the mean-difference extraction produces for those two tasks, it is not a function vector, at least not one that steers. Cross-template transfer of a non-functional vector is not a meaningful object.
The Simpson's paradox
Plot every transfer pair as (cosine, OOD accuracy), colour by task, fit lines.
Aggregate over all three tasks and the slope is steeply negative ( in Base). Toggle "per-task fits" and the negative trend dissolves: within antonym, in Base and in Instruct; within capitalize, in Base; within sentiment-flip, in Base.
The aggregate slope is not telling you about a within-task geometry-vs-behaviour relationship. It is telling you that the broken tasks happen to live in the upper-left of the plane (high cosine, near-zero accuracy) and the working task lives in the centre-right (mid cosine, mid-to-high accuracy). Mixing the two produces a phantom correlation that a per-task analysis never would.
Of the 438 dissociation cases in Base (pairs with and OOD accuracy ), 97.7% come from capitalize and sentiment-flip. In those cases there was no behaviour left for the geometry to predict.
Why broken tasks have artificially high cosine
Capitalize and sentiment-flip share verbatim input words across templates: hello appears in every template-style for capitalize; great movie appears in every template-style for sentiment-flip. The mean-difference at early layers (especially L2) measures the residual-stream response to those shared input tokens, and the FVs end up encoding something about the input wording rather than the task. Two FVs trained on the same word list under different templates will be highly cosine-similar without doing anything functional. The cosine is high and the transfer is zero, but only because there was no transfer to be had.
Antonym, by contrast, draws distinct word pairs under every template (cold/hot, soft/hard, and so on), so the surface-feature contribution averages out and the residual is closer to a task signal.
What RLHF does and does not do
The natural follow-up is whether instruction-tuning fixes the problem.
Absolute dissociation count
Conditional dissociation rate
The absolute count of dissociation cases drops from 438 in Base to 70 in Instruct, about a 6× reduction. But the conditional dissociation rate, given that a pair has , barely moves: 97.7% in Base versus 97.2% in Instruct. Instruct just has fewer high-cosine pairs to begin with (mean cosine 0.435 against 0.603), so fewer clear the threshold. Once a pair is above it, instruction-tuning has not changed what high cosine is worth.
One thing here is a genuine mechanistic difference rather than a distributional one: the layers where patching matters most move from L10 to L18 in Base to L16 to L24 in Instruct, so instruction-tuning does appear to push this computation later.
Per-task alignment varies sharply across layers
The L2/L16/L32 alignment table from the paper:
| Task | L2 mean cos | L16 mean cos | L32 mean cos | IID acc |
|---|---|---|---|---|
| Antonym (Base) | 0.62 | 0.48 | 0.31 | 0.573 |
| Capitalize (Base) | 0.83 | 0.74 | 0.55 | 0.000 |
| Sentiment (Base) | 0.78 | 0.66 | 0.42 | 0.021 |
Capitalize and sentiment-flip have the highest mean inter-template cosine at every layer, and the lowest IID steering accuracy. Cosine and behaviour are pulling in opposite directions across tasks.
The early-layer cosines for the broken tasks are 0.83 and 0.78. The early-layer cosine for the working task is 0.62. Higher cosine, lower accuracy: exactly the inversion that produces the phantom aggregate negative slope.
What I take from it
Two of three tasks failed a prerequisite I had not thought to check, so the cross-template transfer claims I set out to make are off the table. What replaces them is a check to run first: measure IID steering accuracy per task before reading anything into an aggregate across tasks. If extraction works for some and not others, pooling them produces a correlation that exists within none of them.
This is three tasks on one model family, which is not enough to say how common the pattern is. It is enough to say the check is cheap and I should have run it at the start.
The follow-up (Steerable but Not Decodable) scales the design to 12 tasks and 6 models with that check as a hard prerequisite. At that scale the aggregate correlation here does not survive either: pooled r lands between −0.20 and +0.13, and cosine adds almost nothing once you know which task you are looking at.
Setup
- Model: Llama-3.1-8B Base + Instruct, 32 transformer layers, .
- Extraction layers: 16 even layers, .
- Extraction set: 15 input/output pairs per template; FVs are the mean-of-ICL minus mean-of-base residual at the final token.
- Steering scalar: swept over ; best- accuracy is reported per configuration.
- Test set: 50 inputs per template, with a case-insensitive substring criterion for antonym and sentiment-flip and an exact-case criterion for capitalize.
- Pearson r reporting: computed over all (cosine, OOD-accuracy) pairs at the model level for the aggregate and within the task subset for the per-task numbers.
References
- Todd et al. (2024). Function vectors in large language models.
- Hendel et al. (2023). In-context learning creates task vectors.
- Park et al. (2023). The linear representation hypothesis.
- Sclar et al. (2023). Quantifying language models' sensitivity to spurious features in prompt design.