发表机构
Meta AI(Meta AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出循环缩放定律,联合建模循环与稀疏性,预测循环MoE损失,并证明其在万亿token规模下带来参数效率与性能优势。
AI 中文摘要
循环变压器和混合专家(MoE)提供了互补的高效扩展途径:循环在固定参数下增加计算深度,而MoE稀疏性在固定活跃计算下扩大总容量。然而,现有的缩放定律孤立地建模循环或稀疏性。在这项工作中,我们引入了循环缩放定律,这是第一个同时建模循环、稀疏性、模型大小和数据的缩放定律。其核心是一个有界的、稀疏性条件的循环映射,该映射刻画了循环带来的有效参数增益以及稀疏性如何提高这一增益。这些定律比先前的替代方案更准确地预测了循环模型的留出损失,并将标准密集和MoE缩放定律作为特例恢复。除了预测之外,拟合的定律为在计算和内存约束下设计循环MoE模型提供了原则性基础。下游评估进一步证明了这两个轴的互补优势:稀疏性在活跃参数效率上带来约3倍的提升,循环在推理的总参数效率上带来约2倍的提升,而联合缩放进一步推进了性能前沿。作为实际扩展,我们展示了这些增益在万亿token规模下依然成立:在匹配的训练计算下,一个具有定律推导循环的循环MoE在推理基准上匹配一个约2倍大的非循环MoE,同时通过循环实现测试时缩放。
英文摘要
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.
Comments19 pages