Suhail Nadaf

Back to blog

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens

· 10 min read

function-vectorslogit-lenstuned-lenssteeringmechanistic-interpretability

The paper was accepted at the ICML 2026 Workshop on Mechanistic Interpretability.

Give Llama-3.1-8B Base the prompt Japan -> and it will not tell you Tokyo. Add one vector to its residual stream at layer 4 and it will, 88% of the time on the template where this works best.

So the answer is in there somewhere by layer 4. The obvious way to check is the logit lens: take the residual stream at that layer, push it through the model's own unembedding matrix, and read off which tokens it favours. At the layer where the steering works, that readout returns /******/, .MixedReality, and ati. Across the whole task, the correct capital shows up in the lens's top ten tokens 5.6% of the time, and its best layer for that is layer 32, twenty-eight layers after the intervention that made the model right.

I expected the opposite. The study was designed around the hypothesis that steering fails exactly where the information is absent, so the lens and the steering should fail together. They come apart instead, and they come apart in the direction that is harder to explain.

How far the pattern goes

Twelve tasks across five categories: lexical retrieval (antonym, synonym, hypernym), factual retrieval (country-capital, English-Spanish, object-color), morphological transform (past-tense, plural), character and surface (capitalize, first-letter, reverse-word), and compositional (sentiment-flip). Six models from three families, all 7 to 9B: Llama-3.1-8B, Gemma-2-9B and Mistral-7B-v0.3, each as Base and Instruct. Eight templates per task, which gives 56 directed cross-template pairs per task per model and 4,032 in total.

Sorting the 72 task-by-model cells by whether steering works and whether the lens reads the answer: 58 do both, 10 steer without being readable, 3 are readable without steering, and 1 does neither. The gap runs one way. Steering accuracy meets or exceeds logit-lens accuracy in every single cell, and the widest is first-letter on Llama Instruct, where steering hits 0.960 and the lens 0.047.

The pipeline runs eight gated stages, and the readability and patching stages only emit for tasks that pass the in-distribution steering check from the earlier study. Without that check the whole thing measures nothing, which is what the earlier study was about.

baselineno-steering accextractFV per templateprobelinear probessteerα sweep, IID + OODanalyzecosine, transfermechanisticpatching · IID-gatedreadabilitylogit + tuned lensfiguresIID-gated · no learned probeslogit lens (frozen unembed)tuned lens (per-layer affine)

The eight stages. Mechanistic and readability run only on tasks that pass the in-distribution steering check, so the dissociation I report is for tasks the model demonstrably can perform via steering.

The 12-task × 6-model grid

In-distribution steering works for far more tasks than the earlier three-task study suggested.

0.001.00low → high
IID steering accuracy for six models across twelve tasks. Steering reaches 0.88 on country_capital for Llama-3.1-8B Base while the logit lens reads 0.06 there, and 0.96 versus 0.05 on first_letter for Llama-3.1-8B Instruct: the models are steerable on tasks whose answers cannot be decoded from the residual stream. Each cell below is announced with its model, task and value.
model \ task
antonym
synonym
hypernym
country_capital
english_spanish
object_color
past_tense
plural
capitalize
first_letter
reverse_word
sentiment_flip
Llama-3.1-8B (Base)
0.74
0.42
0.09
0.88
0.84
0.36
0.94
0.97
0.62
0.70
0.04
0.05
Llama-3.1-8B (Instruct)
0.86
0.55
0.30
0.95
0.90
0.65
0.98
0.99
0.84
0.96
0.20
0.18
Gemma-2-9B (Base)
0.85
0.62
0.90
0.85
0.80
0.55
0.92
0.96
0.80
0.78
0.39
0.20
Gemma-2-9B (IT)
0.87
0.68
0.78
0.92
0.88
0.72
0.95
0.99
0.83
0.85
0.35
0.16
Mistral-7B (Base)
0.78
0.58
0.45
0.24
0.82
0.40
0.92
0.85
0.55
0.36
0.03
0.10
Mistral-7B (Instruct)
0.82
0.60
0.50
0.30
0.86
0.50
0.94
0.88
0.60
0.50
0.05
0.14
IID steering accuracy and intermediate-layer readability across 12 tasks × 6 models. The "dissociation" view highlights cells where steering succeeds but the answer cannot be read from the residual stream, including country_capital on Llama-3.1-8B Base (steering 0.88, logit-lens 0.06) and first_letter on Llama Instruct (steering 0.96, logit-lens 0.05). Per-category ranges and the named cells are from the study; where a single cell was not extracted directly it is estimated inside its published range, so read the grid for the pattern and the prose for the counts.

On the steering view, morphological transforms (past-tense, plural) clear 0.85 on every model and reach 1.00 on three of six; factual retrieval (country-capital, English-Spanish) clears 0.80 on most models, with Mistral the exception; lexical retrieval (antonym, synonym) clears 0.55 to 0.85; character-level tasks are uneven, with capitalize between 0.55 and 0.87, first-letter between 0.36 and 0.96, and reverse-word failing everywhere. Reverse-word and sentiment-flip do not clear the gate and drop out of the readability analysis.

The logit-lens view does not track any of that. Even the tasks where steering exceeds 0.90 give top-10 readout accuracies between 0.05 and 0.25. The steering − logit-lens diff view shows it directly: 60 of 72 cells have a negative gap, and thirteen of those exceed 0.50.

A worked example

For country_capital on Llama-3.1-8B (Base), a function vector applied at layer 14 steers the model to the correct answer 88 percent of the time, while the correct answer does not appear in the top three logit-lens predictions and does not appear in the top three tuned-lens predictions at the same layer.

Japan -> Tokyo · best layer 14
IID steering
88%
FV at L14 steers the model to the correct answer.
Logit lens · top 3answer absent
  • /******/7.0%
  • .MixedReality4.0%
  • ati3.0%
Tuned lens · top 3answer absent
  • Tok5.0%
  • the4.0%
  • .MixedReality4.0%
Steering succeeds where readout fails. The tuned lens (a learned per-layer affine basis correction) closes the gap on at most one of fourteen steerable-but-not-decodable pairs in the full study; the rest stay invisible to vocabulary-space readers. The steering accuracies are measured; the three readout tokens shown per lens are drawn from the top-50 lists in the appendix, standing in for the full distribution.

The first case is the country-capital one from the opening. The last case is the opposite regime: antonym on Llama-3.1-8B Instruct steers at 86% and "cold" shows up in the top-3 of both lenses. Antonym is the task the linear-representation hypothesis was originally formulated about, and it is the one case here where both lenses read the answer.

Why the logit lens fails

The logit lens projects the residual stream through the frozen unembedding matrix and reports the top tokens. That assumes intermediate representations are a smooth interpolation toward the final layer and that the unembedding is the right basis for reading them. For most of these tasks the second assumption is wrong. Averaged across all six models, the top-50 logit-lens tokens at the steering layer contain a correct-output token for 1.8% of tasks. Even on past-tense and English-Spanish, where steering exceeds 0.90, that fraction sits between 0.000 and 0.048.

First-letter is the only systematic exception, and only in a useless way: 0.84 of its top-50 tokens are single characters. The window fills with the right kind of token and the wrong ones.

The tuned lens controls for basis

The obvious objection is that the unembedding is simply the wrong coordinate frame for an intermediate layer, and a lens trained to correct for that would read the answer fine. The tuned lens does exactly that: a per-layer diagonal affine translator, trained on each model's own activations to map layer ℓ onto the final layer. If the dissociation is a basis artefact, this should dissolve it.

Llama-3.1-8B (Base)

Llama-3.1-8B (Instruct)

Gemma-2-9B (Base)

Gemma-2-9B (IT)

Mistral-7B (Base)

Mistral-7B (Instruct)

logit lens tuned lens
Per-model layerwise top-1 readability, averaged over the working tasks. The tuned lens learns per-layer translators on the model's own outputs and corrects basis-mismatch artifacts; on Mistral the correction is large in magnitude (~25% training improvement) but still leaves the readability curve well below steering accuracy. The per-layer curves are drawn from the per-family means rather than plotted per layer, so the shape is the claim here and the individual layer values are not measurements.

The correction it learns is real and varies a lot by family: Mistral gets about 25% mean improvement in reconstruction, Llama 11%, Gemma-2 under 3%. None of that moves the dissociation. Of fourteen steerable-but-not-decodable pairs, the tuned lens closes one, first-letter on Llama-3.1-8B Instruct, where 0.047 climbs to 0.128 and only just crosses the threshold. For the other thirteen the mean change is −0.005, not distinguishable from zero (p = 0.113). Country-capital on Llama Base does not move at all: 0.056 under both lenses against 0.880 steering.

There is also an anti-correlation I do not have a good account of. The families that need the largest basis correction get the worst readability outcomes overall (r = −0.478, p < 10⁻⁶).

What a stronger decoder finds

A diagonal affine map is still a linear decoder, so the honest next question is whether the answer is there in a form no linear readout can reach. I trained 2-layer MLP probes on the same activations, with a Hewitt and Liang control task to check the probe was reading the model rather than memorising the labels.

Half the dissociation dissolves. Of the ten steerable-not-decodable cells, five close under the MLP probe, which means the answer was present at that layer and nonlinearly encoded. The other five do not: even with the extra capacity, the probe recovers no input-conditional task structure, while steering on those same cells runs between 0.46 and 0.90.

So the strong version of the claim is wrong. In half these cases the information is there and the lens is just too weak an instrument to see it. The claim that survives is narrower: for five cells, every decoder I tried came back empty while the vector still worked. A bigger probe or a different architecture could close some of those too, and the control task cannot rule that out.

What I think is going on

The measured result is the dissociation: steering succeeds at layers where the readouts I tried recover nothing. What follows from that is less certain.

My reading is that the vector does not carry the answer. It sets up a computation, and the remaining twenty-odd layers execute it, which is why nothing intermediate looks like "Tokyo" and only the final projection lands on it. Two other things fit that picture. Style alignment between templates has no effect on transfer (p > 0.39 on every model), and cosine similarity between source and target vectors adds at most ΔR² = 0.011 once you know which task it is. Transfer is predicted by task identity and not by the geometry of the vectors, which is what you would expect if the vector names an operation rather than pointing at an output.

That is an interpretation of a null, and nulls are weak evidence for any particular story about what is there instead. What I would defend is the narrower practical version: a vocabulary-space readout does not characterise a steering direction, and a lens returning nothing is not evidence that the direction is absent.

What this does to the linear-representation hypothesis

Not much, directly. Antonym demonstrates it, and the tuned-lens curves climb monotonically with depth on every model. What the data pushes on is which linear structure is being claimed. Linear structure exists here, but on these tasks a good deal of it lives in directions that reshape the rest of the forward pass rather than in the directions the unembedding reads.

Limitations

Every task here has a single-token or short answer, and open-ended generation may behave differently. All six models are 7 to 9B, and larger models may line their intermediate representations up with the unembedding more closely, which would shrink the gap. I only tested mean-difference extraction, so vectors trained by other methods may look nothing like this. Eight templates per task is more than prior work and still not adversarial coverage. And this is observational on top of behavioural: I measure what the lenses see and what the vectors do, not the circuit the effect travels through.

Reproducibility

The eight stages cache intermediate state by config hash. The mechanistic and readability stages only emit for tasks that clear the 0.10 in-distribution threshold at stage 4. All six tuned lenses were trained from scratch on the same model outputs, with per-layer training curves in the repository. The dissociation table, every heatmap cell and every readability number live in the artefacts directory and reproduce from pipeline.py.

References

  • Todd et al. (2024). Function vectors in large language models.
  • Hendel et al. (2023). In-context learning creates task vectors.
  • Belrose et al. (2023). Eliciting latent predictions from transformers with the tuned lens.
  • Hewitt and Liang (2019). Designing and interpreting probes with control tasks.
  • Park et al. (2023). The linear representation hypothesis.
  • Sclar et al. (2023). Quantifying language models' sensitivity to spurious features in prompt design.