Suhail Nadaf

Back to blog

A Simpson's Paradox Hiding in Function-Vector Cross-Template Transfer

· 10 min read

function-vectorssteeringsimpsons-paradoxmechanistic-interpretabilitynegative-result

I set out to study cross-template function-vector (FV) transfer as a robustness probe. If a vector trained on template A still steers under template B, the underlying representation is template-invariant in a meaningful sense. I extracted 384 FVs from Llama-3.1-8B (Base + Instruct) across three tasks × eight templates × sixteen layers and evaluated all 2,688 cross-template transfer pairs per model. The aggregate cosine-vs-accuracy correlation came out at r=−0.572r = -0.572 in Base and r=−0.361r = -0.361 in Instruct, both with p<10−83p < 10^{-83}. That looked like a clean negative finding. It was a Simpson's paradox.

3 tasks × 8 templates × 16 layersantonymcapitalizesentimenttemplates A B C D E F G Hextractmean-diff at L2…L32per-task: each near-flatantonym r ≈ −0.15capitalize r ≈ −0.10sentiment r ≈ +0.25aggregatemix tasksaggregate: r = −0.57cosine ↑, accuracy ↓

The 384-FV extraction grid (left) is reduced to a single aggregate correlation (right). Within each task the cosine-vs-accuracy slope is near zero or positive; the negative aggregate slope appears only after mixing tasks of very different IID accuracies. That is a Simpson's paradox.

The setup I expected to study

The mean-difference function-vector method (Todd et al., 2024) extracts a single direction FVt,k,ℓ\text{FV}_{t,k,\ell} for task tt at template kk, layer ℓ\ell:

FVt,k,ℓ=E[hℓ(ICL)]−E[hℓ(base)]\text{FV}_{t,k,\ell} = \mathbb{E}\bigl[h_\ell^{(\text{ICL})}\bigr] - \mathbb{E}\bigl[h_\ell^{(\text{base})}\bigr]

The expectation is taken over 15 input/output pairs per template at the final-token position. Steering at inference time adds the FV with a scalar:

hℓ′=hℓ+α⋅FVt,k,ℓ,α∈{0.5,1.0,1.5,2.0,2.5}h_\ell' = h_\ell + \alpha \cdot \text{FV}_{t,k,\ell},\quad \alpha \in \{0.5, 1.0, 1.5, 2.0, 2.5\}

I extract on Llama-3.1-8B Base and Instruct (32 layers, dmodel=4096d_\text{model} = 4096), at 16 even layers ℓ∈{2,4,…,32}\ell \in \{2,4,\ldots,32\}, for three tasks: antonym (semantic retrieval), capitalize (character-level transformation), and sentiment-flip (compositional). Eight templates per task: natural, symbolic, functional, verbose, instruction, question, roleplay, formal. Total: 384 FVs per model.

The headline question is cross-template robustness: extract on template A, evaluate on template B. If FVs encode the task and not the template, transfer accuracy should track cosine alignment between the source and target FVs.

The IID diagnostic that broke the study

Before mixing source-template and target-template, the simplest sanity check is diagonal evaluation: extract on template A, test on template A. If the FV does not steer in-distribution, there is no behaviour to transfer.

TaskBase IID accInstruct IID acc
Antonym0.5730.649
Capitalize0.0000.111
Sentiment-flip0.0210.059

Capitalize FVs steer the Base model to the correct answer 0% of the time. There is no behaviour to transfer.

Capitalize is exactly zero in the Base model. Sentiment-flip is 0.021. Whatever the mean-difference extraction produces for those two tasks, it is not a function vector, at least not one that steers. Cross-template transfer of a non-functional vector is not a meaningful object.

The Simpson's paradox

Plot every transfer pair as (cosine, OOD accuracy), colour by task, fit lines.

Antonym (IID 0.57 / 0.65)Capitalize (IID 0.000 / 0.111)Sentiment-flip (IID 0.021 / 0.059)Aggregate fit r = -0.572
Cosine alignment vs. cross-template OOD transfer accuracy. Bin size encodes count. Toggle "per-task fits" to see the aggregate negative trend dissolve. Illustrative bins reproducing the published correlations.

Aggregate over all three tasks and the slope is steeply negative (r=−0.572r = -0.572 in Base). Toggle "per-task fits" and the negative trend dissolves: within antonym, r=−0.152r = -0.152 in Base and +0.263+0.263 in Instruct; within capitalize, r=−0.099r = -0.099 in Base; within sentiment-flip, r=+0.253r = +0.253 in Base.

The aggregate slope is not telling you about a within-task geometry-vs-behaviour relationship. It is telling you that the broken tasks happen to live in the upper-left of the plane (high cosine, near-zero accuracy) and the working task lives in the centre-right (mid cosine, mid-to-high accuracy). Mixing the two produces a phantom correlation that a per-task analysis never would.

Of the 438 dissociation cases in Base (pairs with cos⁡>0.80\cos > 0.80 and OOD accuracy <0.40< 0.40), 97.7% come from capitalize and sentiment-flip. In those cases there was no behaviour left for the geometry to predict.

Why broken tasks have artificially high cosine

Capitalize and sentiment-flip share verbatim input words across templates: hello appears in every template-style for capitalize; great movie appears in every template-style for sentiment-flip. The mean-difference at early layers (especially L2) measures the residual-stream response to those shared input tokens, and the FVs end up encoding something about the input wording rather than the task. Two FVs trained on the same word list under different templates will be highly cosine-similar without doing anything functional. The cosine is high and the transfer is zero, but only because there was no transfer to be had.

Antonym, by contrast, draws distinct word pairs under every template (cold/hot, soft/hard, and so on), so the surface-feature contribution averages out and the residual is closer to a task signal.

What RLHF does and does not do

The natural follow-up is whether instruction-tuning fixes the problem.

Absolute dissociation count

Pairs with cos > 0.80 and OOD acc < 0.40
438 → 70 (≈6× reduction)

Conditional dissociation rate

Given cos > 0.80, P(OOD acc < 0.40)
97.7% → 97.2% (essentially unchanged)
RLHF reduces the count of high-cosine pairs but not the rate at which such pairs fail to transfer. The mechanism is a distribution shift, not a fix.

The absolute count of dissociation cases drops from 438 in Base to 70 in Instruct, about a 6× reduction. But the conditional dissociation rate, given that a pair has cos⁡>0.80\cos > 0.80, barely moves: 97.7% in Base versus 97.2% in Instruct. Instruct just has fewer high-cosine pairs to begin with (mean cosine 0.435 against 0.603), so fewer clear the threshold. Once a pair is above it, instruction-tuning has not changed what high cosine is worth.

One thing here is a genuine mechanistic difference rather than a distributional one: the layers where patching matters most move from L10 to L18 in Base to L16 to L24 in Instruct, so instruction-tuning does appear to push this computation later.

Per-task alignment varies sharply across layers

The L2/L16/L32 alignment table from the paper:

TaskL2 mean cosL16 mean cosL32 mean cosIID acc
Antonym (Base)0.620.480.310.573
Capitalize (Base)0.830.740.550.000
Sentiment (Base)0.780.660.420.021

Capitalize and sentiment-flip have the highest mean inter-template cosine at every layer, and the lowest IID steering accuracy. Cosine and behaviour are pulling in opposite directions across tasks.

The early-layer cosines for the broken tasks are 0.83 and 0.78. The early-layer cosine for the working task is 0.62. Higher cosine, lower accuracy: exactly the inversion that produces the phantom aggregate negative slope.

What I take from it

Two of three tasks failed a prerequisite I had not thought to check, so the cross-template transfer claims I set out to make are off the table. What replaces them is a check to run first: measure IID steering accuracy per task before reading anything into an aggregate across tasks. If extraction works for some and not others, pooling them produces a correlation that exists within none of them.

This is three tasks on one model family, which is not enough to say how common the pattern is. It is enough to say the check is cheap and I should have run it at the start.

The follow-up (Steerable but Not Decodable) scales the design to 12 tasks and 6 models with that check as a hard prerequisite. At that scale the aggregate correlation here does not survive either: pooled r lands between −0.20 and +0.13, and cosine adds almost nothing once you know which task you are looking at.

Setup

  • Model: Llama-3.1-8B Base + Instruct, 32 transformer layers, dmodel=4096d_\text{model} = 4096.
  • Extraction layers: 16 even layers, ℓ∈{2,4,…,32}\ell \in \{2,4,\ldots,32\}.
  • Extraction set: 15 input/output pairs per template; FVs are the mean-of-ICL minus mean-of-base residual at the final token.
  • Steering scalar: α\alpha swept over {0.5,1.0,1.5,2.0,2.5}\{0.5, 1.0, 1.5, 2.0, 2.5\}; best-α\alpha accuracy is reported per configuration.
  • Test set: 50 inputs per template, with a case-insensitive substring criterion for antonym and sentiment-flip and an exact-case criterion for capitalize.
  • Pearson r reporting: computed over all (cosine, OOD-accuracy) pairs at the model level for the aggregate and within the task subset for the per-task numbers.

References

  • Todd et al. (2024). Function vectors in large language models.
  • Hendel et al. (2023). In-context learning creates task vectors.
  • Park et al. (2023). The linear representation hypothesis.
  • Sclar et al. (2023). Quantifying language models' sensitivity to spurious features in prompt design.