← Architecture & Pretraining · Learning notes

Why does the first token after switch-off not prove persistence?

In my steering experiments, the first generated token seemed to retain the intervention. The timing turned out to be crucial.

The intuition

The logits at the last prompt position predict the first new token. If steering was active over that prompt position, that prediction still came from steered computation. The next unsteered forward position predicts the second generated token.

last prompt position -> token 1
token 1 position -> token 2

There is another complication: the first selected token can itself change the future context. Fixed-token scoring and free generation therefore answer different questions. This alignment check came from the experiment; the paper below is background on activation steering.

Related reading · Steering Language Models With Activation Engineering