← Architecture & Pretraining · Learning notes

Why does SwiGLU need three matrices?

The feed-forward block suddenly had two input projections instead of one. I wanted to know what the extra branch was buying.

The intuition

Both branches expand D features to F features. One passes through SiLU and gates the other by elementwise multiplication. A third projection maps the resulting F features back to D.

y = W2(SiLU(W1(x)) * W3(x))

The gate is a continuous feature-wise modulation, not a choice of tokens or a miniature attention layer. Thinking of one branch as content and the other as its modulation makes the shape bookkeeping natural.

Related reading · GLU Variants Improve Transformer