Is a hidden state really richer than the logits?
In our discussion of latent communication, I objected that the final hidden state was only one matrix multiplication away from vocabulary scores.
The intuition
That objection matters. If the output matrix has full column rank, exact logits can mathematically determine its input vector. The obvious compression occurs when we sample or take argmax and retain only one token.
h in R^D -> W_out h in R^V -> one tokenThe useful comparison is not simply hidden state versus logits. It is continuous representation versus a selected token, together with the receiving model’s interface, precision, and training objective. A vector is not automatically a useful message.
Related reading · Using the Output Embedding to Improve Language Models