Why does SwiGLU need three matrices?
The feed-forward block suddenly had two input projections instead of one. I wanted to know what the extra branch was buying.
The intuition
Both branches expand D features to F features. One passes through SiLU and gates the other by elementwise multiplication. A third projection maps the resulting F features back to D.
y = W2(SiLU(W1(x)) * W3(x))The gate is a continuous feature-wise modulation, not a choice of tokens or a miniature attention layer. Thinking of one branch as content and the other as its modulation makes the shape bookkeeping natural.
Related reading · GLU Variants Improve Transformer