← Architecture & Pretraining · Learning notes

Are Adam’s bias correction and normalization the same thing?

I understood the recursive compression of past gradients. The initialization correction and the second-moment denominator were the parts I kept mixing up.

The intuition

The moving averages start at zero, so their early values are biased toward zero. Dividing by 1 minus beta to the power t corrects that effect. With a constant gradient of 2 and beta1=0.9, the first moment is initially 0.2; correction brings it to 2.

m_hat = m / (1 - beta1**t)
step = lr * m_hat / (sqrt(v_hat) + eps)

Dividing the corrected first moment by the square root of the corrected second moment is a separate operation: it scales each coordinate using its own gradient history. It does not normalize the entire gradient vector to unit length.

Related reading · Adam: A Method for Stochastic Optimization