Weekly Paper Notes — one of the top picks from the 2026-10-03 CS paper digest. Area: Distributed Computing.

Authors: Stuart H. Sul, Nash Brown (Stanford, Cursor Research), Henry Wildermuth, William Lin, Federico Cassano (Cursor Research), Christopher Ré (Stanford) arXiv: 2609.36070 · PDF · Code

TL;DR

Mixture-of-Kittens (MoK) is the MoE training kernel Cursor uses to train Composer on GB300 NVL72 racks. The paper’s opening finding is uncomfortable for anyone who has invested in MoE communication libraries: on a 72-GPU NVLink domain, existing expert-parallel systems built for InfiniBand often lose to a naive PyTorch + NCCL baseline, and in nearly half of the evaluated configurations the naive baseline beats every alternative.

The authors trace that to three assumptions that stop holding once all expert-parallel ranks sit on one single-hop fabric. Pull-based transfers are now as cheap as push-based ones, so communication direction can be chosen per operator to keep completion signals local. The right overlap granularity varies from 512 to 32,768 tokens depending on the workload, so it has to be a tunable rather than a design constant. And the Grace CPUs on NVL72 are slow enough at dispatching PyTorch operations (1.46–2.97x slower than a DGX B300 host) that any CPU-GPU synchronization becomes expensive, so buffer management moves entirely onto the device.

All of that lands in one persistent, bitwise-deterministic megakernel that fuses dispatch, shared and routed expert FFNs, gated activation, and combine. Against the strongest public baseline it reaches up to 2.37x on MXFP8 forward and 1.78x on MXFP8 backward at EP=64. Swapped into Cursor’s production stack on 512 GB300 GPUs, end-to-end training throughput goes from 761 to 1,070 tokens/s/GPU.

What problem is the paper actually attacking?

Expert parallelism shards routed experts across GPUs, so every MoE layer does two all-to-alls: dispatch sends each token to the ranks holding its top-k experts, and combine brings the expert outputs back for the router-weighted sum. At scale those two exchanges plus the expert GEMMs eat more than half of step time, which is why there’s a thick literature on overlapping them. GShard and Switch Transformer established the pattern and accepted token dropping to keep buffers fixed-size. MegaBlocks removed the dropping with block-sparse kernels but needed the host to learn per-expert counts before allocating. Tutel pipelined all-to-all with expert compute adaptively. DeepEP gave DeepSeek-V3 dedicated dispatch/combine kernels for coarse-grained overlap across nodes. Comet and FlashMoE pushed overlap down to tile granularity with device-initiated communication. MegaMoE in DeepGEMM fused the whole thing into a single kernel, but only for DeepSeek-V4 inference in FP8×FP4.

Nearly all of that work assumed a scale-out network: a handful of GPUs on NVLink, then a slower RDMA fabric. In that world push beats pull for fine-grained RDMA, coarse chunks amortize network latency, and host round-trips are tolerable because a fast x86 CPU is driving them.

NVL72 changes the cost model. Seventy-two Blackwell GPUs share a non-blocking, single-hop NVLink fabric at 900 GB/s unidirectional per GPU, so an EP group of 64 fits inside one domain, and the interconnect starts to behave like one big memory system rather than a network. Nvidia’s roadmap goes to NVL144, 576 and 1152, so this isn’t a one-off.

The observation that motivates MoK is the benchmark result above. If a straightforward NCCL all-to-all can beat purpose-built MoE communication libraries, the libraries are encoding costs that no longer exist on this hardware. MoK starts from a cost model rather than from an existing kernel:

C = max(T_comp, T_mem, T_comm) + T_non-overlap + T_sync + T_host

Perfect overlap makes the first term a maximum. Everything else is overhead to be removed: the unoverlapped first dispatch and last combine, the signals and fences between workers, and host-side launches and synchronizations.

The mechanism: three changes, one kernel

MoK overview. Communication SMs pull tokens from peer GPUs into a fixed-size macrobatch ring buffer; compute SMs run each minibatch through the expert FFN; communication SMs push outputs back. Router projection, top-k, and dispatch scheduling run on the GPU before launch. Figure: MoK’s SM partitioning, minibatch pipeline, and pre-launch GPU-side scheduling. Source: Sul et al., arXiv:2609.36070, Figure 1.

Pull dispatch, push combine

Under fine-grained overlap, T_sync dominates because every tensor-core task must know its input has arrived. With push-based dispatch, every peer sends a completion signal for every chunk it writes, and the receiving GPU waits on all of them. With pull-based dispatch, the receiving GPU’s own communication SMs issue the loads and only need to tell local compute SMs when they’re done.

The paper measures this on NVL72 with 7,168 BF16 tokens per GPU: push signalling overhead grows from 39 µs at EP=4 to 135 µs at EP=64 (8–18% of dispatch latency), while pull signalling stays within 4 µs at every EP degree. MoK’s rule is that the GPU receiving the signal should be the one issuing the remote memory operations. Forward is pull dispatch and push combine; backward mirrors it with pull reverse-combine and push reverse-dispatch. The compute/communication handoff never crosses a GPU boundary.

Direction also decides who builds the transfer schedule. MoK runs a separate kernel that produces a two-column table, {src_rank, src_index}, indexed by destination buffer position. It sorts routes by (local expert, per-source ordinal, source rank): the expert key gives each expert a contiguous buffer region for its GEMM, and the (ordinal, rank) key interleaves source ranks round-robin so all NVLink lanes carry similar load. Building it costs 2.1–7.7% of MoE runtime, 1.4–2.2x less than Comet’s equivalent. Push combine reuses the same table by reading the columns as {dst_rank, dst_index}, so all four communication operators share it.

Granularity as a tunable

Call the overlap chunk a minibatch of size b. Small b never saturates tensor cores and pays per-chunk synchronization; large b leaves the first dispatch and last combine exposed. The paper writes the two penalties as P_fine ≈ ⌈N_routed / b⌉ · C_ineff(b, w) and P_coarse ≈ T_dispatch(b) + T_combine(b) and argues b should minimize their sum per workload. Empirically the optimum ranges from 512 to 32,768 tokens, and for a fixed model and token count the best b gives 1.38–3.53x the throughput of the worst. Comet sits at one extreme (an MMA tile), DeepEP’s usual configuration at the other.

Making b cheap to vary takes two things. Synchronization is maximally deferred: communication walks minibatches breadth-first, dispatching everything as fast as it can, while compute walks depth-first, carrying one minibatch through the full FFN before taking the next. They meet only where compute waits for a minibatch to land and combine waits for its outputs. And everything is fused into one persistent kernel with a fixed split between compute SMs and communication SMs that signal through a GPU-local barrier. Without fusion, each minibatch would need several launches across streams, tens of microseconds of T_host apiece, and the hardware scheduler wouldn’t guarantee the two sides actually run concurrently.

A ring buffer instead of a host round-trip

Per-expert token counts are only known after routing. Exact-size allocation needs those counts on the CPU (a sync), and fixed-capacity buffers drop tokens (a quality hit). MoK does neither. It cycles tokens through a device-resident ring buffer, the macrobatch, with dispatch writing one slot while FFNs consume another and combine drains a third. Against an over-allocated buffer that holds every routed token, the ring adds 1.7% latency in geometric mean.

The ring overwrites activations, which matters for training. MoK fuses a ring-aware forward replay into the backward megakernel that recomputes only what the backward pass needs and overlaps it with backward work. It also runs the forward pass over macrobatches in reverse order, so the ring ends full and the backward replay is as short as possible.

The production details

Several features in §3.4 are what make this a training system rather than a benchmark kernel.

Determinism is total: the same inputs give bitwise-identical outputs. Atomics to a shared address are either forced into a fixed order by software scheduling or replaced with writes to a scratch buffer that one worker reduces in a fixed order. The authors cite training ablations and on-policy RL as the reasons, and for RL in particular a nondeterministic MoE layer makes trainer/sampler logprob mismatches hard to tell apart from bugs.

Tasks are scheduled through Blackwell’s Cluster Launch Control, a hardware work-stealing mechanism that lets the persistent kernel yield SMs to a higher-priority stream. That’s how FSDP all-gathers across racks get SMs while the megakernel is resident instead of queuing behind it.

MXFP8 is supported with activation quantization fused into dispatch, the grouped GEMMs, and the gated activation. The shared expert stays in BF16 for stability and is scheduled into the first dispatch window, when compute SMs would otherwise be idle. Router weight gradients use SonicMoE’s trick of taking the inner product of the gated activation with the down-projection input gradient, so the down-projection output is never materialized.

Results

MoE layer throughput per GPU on a GB300 NVL72 for four model shapes, forward and backward, BF16 and MXFP8, EP=64 with 2,048 tokens per GPU. MoK leads every configuration. Figure: Effective expert-FFN TFLOP/s per GPU at EP=64, 2,048 tokens/GPU. Source: Sul et al., arXiv:2609.36070, Figure 7.

The layer benchmarks use the MoE shapes of Kimi K2.7 Code, GLM-5.2, Qwen3.5-397B-A17B and DeepSeek-V4-Pro, and compare against every public implementation that does both forward and backward on NVL72: NCCL + PyTorch, DeepEP + PyTorch, DeepEP + TransformerEngine, and HybridEP + Megatron. Timing takes the slowest rank’s latency, median of 100 runs after 500 warmups, and includes schedule construction.

At EP=64 and 2,048 tokens/GPU, MoK’s best speedups over the fastest baseline are 2.37x (MXFP8 forward), 1.78x (MXFP8 backward), 1.92x (BF16 forward) and 1.58x (BF16 backward). The ordering follows from the arithmetic. BF16 peak is half of MXFP8 on GB300, and backward does twice the FLOPs of forward, so both push the layer toward compute-bound and leave less communication for MoK to hide.

The sensitivity table is more telling than the headline. Across EP=16 and EP=64 and 1,024 to 4,096 local tokens, forward speedups run from 1.29x to 2.74x, and they’re largest at the smallest token count, where scheduling and communication overheads are the biggest share. In most of the EP=16 rows the best baseline is plain NCCL + PyTorch, which supports the opening claim. The authors point out that the small-token regime is where large training runs end up, because global batch size can’t grow indefinitely, so adding GPUs means fewer tokens per GPU.

End to end, at EP=32 with MXFP8 routed experts across multiple GB300 racks, replacing Cursor’s previous DeepEP-based MoE path with MoK takes throughput from 761.0 to 1,070.2 tokens/s/GPU, a 41% gain. Cursor’s real runs use more GPUs and fewer tokens per GPU than this experiment, which by the same argument should widen the gap.

Why this matters

The useful lesson here is about cost models more than about MoE. Push-based dispatch, coarse chunking, and host-side allocation were each the right choice when they were made, for a fabric where remote reads were expensive and the host was fast. NVL72 inverts both, and the libraries kept the old choices. The paper’s three changes are what you get by re-deriving the design from C = max(...) + overheads on the new hardware. That’s a method other teams can apply when moving to Helios, TPU pods, or Trn2 UltraServers.

The authors also make a forward-looking argument. Vera Rubin raises FP8 throughput 3.5x over GB300 but scale-up bandwidth only 1.67x, so MoE layers get more communication-bound and overlap matters more. Teams planning around the next rack generation should read the paper with that in mind.

It’s a production artifact: open source, the MoE backend in Nvidia NeMo AutoModel, and used to train Composer across tens of thousands of GPUs. Determinism and FSDP coexistence often get skipped in research kernels, and here they were built in from the start.

Some limits are worth keeping in view. All the comparisons are against public baselines. NVL72-native internal stacks at other labs aren’t represented, and the 41% end-to-end number is relative to Cursor’s own previous implementation. Router inputs in the layer benchmarks are standard-normal, which gives far more uniform expert load than real training. Skewed routing, where a few hot experts create stragglers, is what this week’s MegaFlux (arXiv:2610.00671) targets with runtime expert replication, and combining that with MoK’s direction and granularity choices is an obvious follow-up.

Read alongside

  • Lepikhin et al., GShard (2020), and Fedus, Zoph, Shazeer, Switch Transformers (JMLR 2022): the original expert-parallel design and the token-dropping trade-off MoK avoids.
  • Gale, Narayanan, Young, Zaharia, MegaBlocks (MLSys 2023): dropless MoE, and the CPU-sync allocation cost MoK’s ring buffer removes.
  • DeepSeek-AI, DeepSeek-V3 Technical Report (2024), with the DeepEP library: the coarse-grained push-based baseline.
  • Zhang et al., Comet (MLSys 2025), and Aimuyo, Oh, Singh, FlashMoE (NeurIPS 2025): fine-grained, device-initiated overlap at the other end of the granularity range.
  • Guo, Mishra, Cheng, Stoica, Dao, SonicMoE (ICLR 2026): the IO-aware expert FFN techniques MoK reuses, including the router-gradient fusion.
  • Cheng et al., MPK: Mega-Kernelizing Tensor Programs (OSDI 2026): the general megakernel approach MoK builds on.

📄 arXiv abstract · 📄 PDF · 💾 cursor/mixture-of-kittens · 📝 Cursor blog post


Part of the Weekly CS Paper Digest series. Summary written from a close read of the preprint; figures cropped from the arXiv PDF and reproduced here under fair use for educational commentary.