Why can a stable softmax still produce log(0)?
I had already subtracted the maximum. Why was taking the logarithm still unsafe?
The intuition
Subtracting the maximum prevents exponential overflow; a sufficiently small probability can still underflow to zero. For logits [0, -1000], the second log probability should be approximately -1000, even if the probability itself cannot be represented.
loss = torch.logsumexp(logits, -1) - target_logitStay in log space. Computing logsumexp minus the target logit avoids materializing a tiny probability only to take its logarithm. The order of operations matters even when two formulas are algebraically identical.
Related reading · PyTorch: log_softmax