Why are there still two sums in cross-entropy?
One-hot targets collapse the sum over classes. So why did my implementation still need a vocabulary reduction and a mean over positions?
The intuition
The target selects one log probability, but that probability still contains the softmax normalizer over all vocabulary candidates. After computing one loss per position, training averages those losses over the batch and sequence.
loss_at_position = logsumexp(logits) - logits[target]These reductions answer different questions: how much probability belongs to the correct token, and how to aggregate prediction errors. Keeping the vocabulary axis separate from the position axes prevents a surprising number of bugs.
Related reading · CS336: Language Modeling from Scratch