Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
LOOM scales looped Mixture-of-Experts Transformers to 9-12 loops, cutting 1.7B model perplexity from 9.62 to 7.77 with accuracy gains.
The paper diagnoses two obstacles to scaling looped MoE LLMs: hidden-state variance growth that destabilizes deep recurrence, and expert selection collapse where routers repeat the same experts across loops. The proposed LOOM method bounds variance via scaled residual updates, re-injects input embeddings each loop, and uses per-loop routers plus a Looping Residual for computational diversity. Experiments on 100M-1.7B models show stable scaling to 9-12 loops; near-iso-FLOP the 700M model at 5 loops reduces perplexity from 18.36 to 16.54 and improves zero-shot accuracy from 38.84% to 39.53%. Without FLOP matching, a 1.7B model trained on 60B tokens peaks at 9 loops, cutting perplexity from 9.62 to 7.77 and raising zero-shot accuracy from 42.4% to 47.7%. Code is released on GitHub.
- Identifies curse of depth and expert selection collapse as looped-MoE scaling blockers
- Per-loop routers and input re-injection stabilize recurrence to 9-12 loops
- 1.7B model gains 5.2 points zero-shot accuracy and 1.85 perplexity at 9 loops
Full article273 words · extracted from huggingface.co · click to collapse
Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. Second, looped MoE suffers from expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity. Guided by this diagnosis, we propose LOOM, built on a single principle: each loop should contribute new computation while keeping the recurrent state stable. LOOM stabilizes recurrence by scaling residual updates to bound variance growth and re-injecting the input embedding at every loop, and diversifies it through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward. Experiments across 100M-1.7B models show stable scaling to 9-12 loops. Under near-iso-FLOP, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53% over the non-looped baseline. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%. Code is available https://github.com/hed-ucas/LOOM.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.01153