AI 中文总结
针对循环MoE扩展受限的问题,提出LOOM方法,通过稳定循环状态和多样化专家选择,实现9-12次循环,显著提升困惑度和零样本准确率。
AI 中文摘要
循环Transformer为大型语言模型引入了循环深度作为新的扩展轴:通过重复应用共享的Transformer块,它们在不增加参数数量的情况下增加了有效深度。然而,在FLOPs匹配的比较下,循环对于大型MoE LLM的益处仍不明确。主要原因是额外迭代带来的收益迅速减少,甚至可能转为性能下降,因此花费在循环上的额外FLOPs仅带来微小的实质性改进。因此,先前的工作通常只采用两次循环。我们识别出扩展循环MoE的两个主要障碍。首先,循环继承并放大了深度诅咒:随着残差更新的累积,隐藏状态方差随每次迭代增长,这破坏了深层循环的稳定性并导致表示漂移。其次,循环MoE遭受专家选择崩溃:路由器在循环中反复选择相同的专家,因此额外迭代增加了计算量却没有增加计算多样性。基于这一诊断,我们提出了LOOM,它基于一个单一原则:每次循环应贡献新的计算,同时保持循环状态稳定。LOOM通过缩放残差更新以限制方差增长并在每次循环时重新注入输入嵌入来稳定循环,并通过每循环路由器参与不同专家以及循环残差将早期输出向前传递来实现多样化。在100M-1.7B模型上的实验表明,可稳定扩展到9-12次循环。在近似等FLOP条件下,700M模型在5次循环时表现最佳,将困惑度从18.36降至16.54,并将平均零样本准确率从38.84%提升至39.53%,优于非循环基线。在不进行FLOP匹配的情况下,在60B令牌上训练的1.7B模型在9次循环时达到峰值,将困惑度从9.62降至7.77,并将平均零样本准确率从42.4%提升至47.7%。代码可在以下网址获取:https URL。
英文摘要
Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. Second, looped MoE suffers from expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity. Guided by this diagnosis, we propose LOOM, built on a single principle: each loop should contribute new computation while keeping the recurrent state stable. LOOM stabilizes recurrence by scaling residual updates to bound variance growth and re-injecting the input embedding at every loop, and diversifies it through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward. Experiments across 100M-1.7B models show stable scaling to 9-12 loops. Under near-iso-FLOP, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53% over the non-looped baseline. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%. Code is available https://github.com/hed-ucas/LOOM.