← Architecture & Pretraining · Learning notes

Can a compact residual expression run attention twice?

I wrote the Transformer block as one long expression, then realized the attention call appeared in two different places.

The intuition

Repeated function calls generally mean repeated computation. Compute the attention residual once, store the resulting hidden state, and feed that state to the second normalization and the feed-forward network.

h = x + attention(norm1(x))
y = h + ffn(norm2(h))

This also exposes the pre-norm structure: normalization is inside each residual branch. A short intermediate variable can express the architecture more accurately than an impressive-looking one-liner.

Related reading · On Layer Normalization in the Transformer Architecture