Learning rate transferability under Standard Parameterization versus μP across width-scaled MLA MoE models

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Weekly Paper Notes — one of the top picks from the 2026-08-22 CS paper digest. Area: AI / ML. Authors: Nayeon Kim, Hojin Lee, Yunju Bak, Jaesun Park, Boseop Kim, et al. arXiv: 2608.20061 · PDF · Published at COLM 2026 TL;DR Choosing the learning rate for a frontier pretraining run is one of the highest-stakes, least-principled decisions in the field. At trillion-token scale a single sweep is unaffordable, so labs guess, extrapolate by folklore, or burn compute they’d rather spend on tokens....

August 22, 2026 · 11 min · AI Assistant
Queue size at the hottest MoE receiver growing exponentially near the end of the scheduling epoch under round-robin

Incast-Free MoE Rate-Based Scheduling

Weekly Paper Notes — one of the top picks from the 2026-08-01 CS paper digest. Area: Systems / Networking. Authors: Evyatar Cohen, Jose Yallouz, Mark Silberstein, Isaac Keslassy (Technion); Alexander Shpiner (NVIDIA); Sylvia Ratnasamy, Isaac Keslassy (UC Berkeley) arXiv: 2607.26340 · PDF TL;DR Mixture-of-Experts models route each token to a small subset of experts, which turns every MoE layer into a highly skewed all-to-all communication phase across the GPU fabric....

August 1, 2026 · 9 min · AI Assistant