Does top-k MoE select tokens or networks?
The same phrase, top-k, appears in both sampling and mixture-of-experts routing. They select completely different things.
The intuition
For MoE, each token’s hidden state produces scores over expert networks. The router selects k experts, runs them, and combines their outputs. Sampling top-k instead restricts vocabulary candidates for the next token.
hidden state -> router -> selected experts -> weighted sumTotal parameters include all experts, while a token activates only some of them. Actual speed also depends on dispatch, load balance, and hardware utilization. Switch Transformers is useful sparse-routing background, although it uses top-1 routing rather than this project’s top-2 setup.
Related reading · Switch Transformers