Architecture & Pretraining
I trained a 2-layer, 58,784-parameter Transformer prototype with top-2 MoE, pre-norm blocks, and shared input embeddings and output projection weights on a synthetic token sequence. GitHub
Some Interesting Questions
Based on discussions with friends and notes organized with Codex and Claude Code.
- Is an embedding layer doing any computation?
- Why does Linear(3, 2) store a (2, 3) matrix?
- Why does SwiGLU need three matrices?
- Why are there still two sums in cross-entropy?
- Why is logits[:, targets] the wrong selection?
- What is keepdim actually keeping?
- Why can a stable softmax still produce log(0)?
- Why is variance the wrong denominator for RMSNorm?
- Are Adam’s bias correction and normalization the same thing?
- Can a compact residual expression run attention twice?
- Is a hidden state really richer than the logits?
- Does seeing the whole prefix make the last state a summary?
- Why does the first token after switch-off not prove persistence?
- Does a KV cache belong to the input or the model?
- What do tied embeddings actually share?
- Does top-k MoE select tokens or networks?