Assignment 1

Architecture & Pretraining

I trained a 2-layer, 58,784-parameter Transformer prototype with top-2 MoE, pre-norm blocks, and shared input embeddings and output projection weights on a synthetic token sequence. GitHub

Some Interesting Questions

Based on discussions with friends and notes organized with Codex and Claude Code.

← All tracks