The Slowest Kid Problem: How a Super Captain Solves MoE's Biggest Headache
Why a mixture-of-experts layer runs at the speed of its busiest expert, how to measure it, and when predictive prefetching actually helps.
Mixture of Experts02
Subject · 3 posts
Why a mixture-of-experts layer runs at the speed of its busiest expert, how to measure it, and when predictive prefetching actually helps.
Mixture of Experts02
Scalars, vectors, matrices and tensors, the five operations every neural network is made of, and the shape rules that cause most of the bugs.
AI Foundations01
Attention and the transformer block from first principles, plus what changed by 2026: pre-norm, RMSNorm, RoPE, grouped-query attention and the KV cache.