MoK overview: communication SMs pull tokens into a macrobatch ring buffer, compute SMs run expert FFNs, outputs are pushed back

Mixture-of-Kittens: MoE Megakernel for NVL72s

Weekly Paper Notes — one of the top picks from the 2026-10-03 CS paper digest. Area: Distributed Computing. Authors: Stuart H. Sul, Nash Brown (Stanford, Cursor Research), Henry Wildermuth, William Lin, Federico Cassano (Cursor Research), Christopher Ré (Stanford) arXiv: 2609.36070 · PDF · Code TL;DR Mixture-of-Kittens (MoK) is the MoE training kernel Cursor uses to train Composer on GB300 NVL72 racks. The paper’s opening finding is uncomfortable for anyone who has invested in MoE communication libraries: on a 72-GPU NVLink domain, existing expert-parallel systems built for InfiniBand often lose to a naive PyTorch + NCCL baseline, and in nearly half of the evaluated configurations the naive baseline beats every alternative....

October 3, 2026 · 11 min · AI Assistant