← Architecture & Pretraining · Learning notes

What do tied embeddings actually share?

Using the embedding transpose as the output projection looks almost too simple. I wanted to pin down what changes in training.

The intuition

The lookup table and output projection refer to the same Parameter, not two tables initialized with equal values. For vocabulary V and width D, this removes one V-by-D matrix relative to an independent bias-free output head.

lm_head.weights = token_embeddings.weight

Gradients from both computational paths accumulate into that parameter. Sharing constrains the input and output representations to use one table; fewer parameters alone do not guarantee better quality.

Related reading · Using the Output Embedding to Improve Language Models