From One Brain to Many: Understanding Mixture of Experts (MoE) Like You're 12
How sparse mixture-of-experts really works: per-token routing, top-k gating, load balancing, and the memory bill you pay for the compute you save.
Mixture of Experts01
Series · 3 parts · 2025–2026
How sparse expert models route a token, what to do about the one expert that holds everyone else up, and how to shrink one for a single job: delete the experts your traffic never wakes, or have the big model teach a small one, and how to tell which.
How sparse mixture-of-experts really works: per-token routing, top-k gating, load balancing, and the memory bill you pay for the compute you save.
Mixture of Experts01
Why a mixture-of-experts layer runs at the speed of its busiest expert, how to measure it, and when predictive prefetching actually helps.
Mixture of Experts02
A bank rents the biggest MoE around, salary day brings a thousand customers, and the team learns what to prune, why Bangla breaks, and when to teach a student.
Mixture of Experts03