What do tied embeddings actually share?
Using the embedding transpose as the output projection looks almost too simple. I wanted to pin down what changes in training.
The intuition
The lookup table and output projection refer to the same Parameter, not two tables initialized with equal values. For vocabulary V and width D, this removes one V-by-D matrix relative to an independent bias-free output head.
lm_head.weights = token_embeddings.weightGradients from both computational paths accumulate into that parameter. Sharing constrains the input and output representations to use one table; fewer parameters alone do not guarantee better quality.
Related reading · Using the Output Embedding to Improve Language Models