Does a KV cache belong to the input or the model?
My adapter experiments kept running into the same question: what exactly becomes stale when the weights change?
The intuition
Both determine the cache. Keys and values are projections of hidden states, and those hidden states depend on the prefix and earlier model computations. Keeping the tokens fixed does not guarantee the same cached tensors under different parameters.
K_l = H_l W_K,l
V_l = H_l W_V,lThe useful experiment crosses the current reader with the model that wrote the cache. It separates present parameter effects from historical computation. LoRA provides background on parameter adaptation; it does not itself establish that an old cache can be reused safely.
Related reading · LoRA: Low-Rank Adaptation of Large Language Models