← Architecture & Pretraining · Learning notes

Does top-k MoE select tokens or networks?

The same phrase, top-k, appears in both sampling and mixture-of-experts routing. They select completely different things.

The intuition

For MoE, each token’s hidden state produces scores over expert networks. The router selects k experts, runs them, and combines their outputs. Sampling top-k instead restricts vocabulary candidates for the next token.

hidden state -> router -> selected experts -> weighted sum

Total parameters include all experts, while a token activates only some of them. Actual speed also depends on dispatch, load balance, and hardware utilization. Switch Transformers is useful sparse-routing background, although it uses top-1 routing rather than this project’s top-2 setup.

Related reading · Switch Transformers