← Architecture & Pretraining · Learning notes

Why can a stable softmax still produce log(0)?

I had already subtracted the maximum. Why was taking the logarithm still unsafe?

The intuition

Subtracting the maximum prevents exponential overflow; a sufficiently small probability can still underflow to zero. For logits [0, -1000], the second log probability should be approximately -1000, even if the probability itself cannot be represented.

loss = torch.logsumexp(logits, -1) - target_logit

Stay in log space. Computing logsumexp minus the target logit avoids materializing a tiny probability only to take its logarithm. The order of operations matters even when two formulas are algebraically identical.

Related reading · PyTorch: log_softmax