arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超参数缩放定律跨越MoE稀疏性

Hyperparameter Scaling Laws Across MoE Sparsity

Changxin Tian, Kunlong Chen, Jia Liu, Ziqi Liu, Zhiqiang Zhang, Jun Zhou

arXiv 2609.08690首次发表:更新:

发表机构

Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对超稀疏MoE,通过大规模实验揭示最优学习率与批量大小随激活比率变化的规律,提出统一超参数缩放定律,实现跨稀疏性和模型规模的可靠迁移。

AI 中文摘要

混合专家(MoE)模型在不按比例增加训练计算量的情况下扩展模型容量,但增加稀疏性使得可靠的超参数迁移变得具有挑战性。在这项工作中,我们表明传统的超参数缩放定律对于超稀疏MoE是不充分的:最优学习率和批量大小随激活比率变化,而这些变化不能仅由总参数或激活参数数量解释。为了刻画这种依赖性,我们进行了1,800次预训练运行,涵盖六个激活参数规模,模型总非嵌入参数高达6B,处理约20万亿个令牌,成本相当于200,000个H800 GPU小时。我们的结果通过揭示两种缩放机制调和了先前工作中的矛盾发现。在固定稀疏性下,最优批量大小与训练令牌数$D$呈幂律关系,而最优学习率随训练计算量$C$缩放,并且对模型大小和数据之间的分配保持稳健。在跨稀疏性水平上,激活比率$A$作为额外的乘法幂律因子进入这两个关系。这些观察导致了统一的超参数缩放定律,可跨MoE稀疏性水平迁移。大规模评估表明,该缩放形式优于其他函数形式。在一个留出的超稀疏MoE上,该模型总参数为12B,仅激活其专家的1/64,预测的超参数仍接近观测最优值,支持跨模型规模和稀疏性的联合外推。进一步实验证明了跨专家粒度的迁移,并隔离了激活比率与总专家数量的影响。

英文摘要

Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyperparameter transfer challenging. In this work, we show that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone. To characterize this dependence, we conduct 1,800 pre-training runs spanning six activated-parameter scales and models with up to 6B total non-embedding parameters, processing approximately 20 trillion tokens at a cost of 200,000 equivalent H800 GPU-hours. Our results reconcile conflicting findings in prior work by revealing two scaling regimes. At fixed sparsity, the optimal batch size follows a power-law relationship with training tokens $D$, whereas the optimal learning rate scales with training compute $C$ and remains robust to the allocation between model size and data. Across sparsity levels, the activation ratio $A$ enters both relationships as an additional multiplicative power-law factor. These observations lead to unified hyperparameter scaling laws that transfer across MoE sparsity levels. Large-scale evaluation shows that the scaling form outperforms alternative functional forms. On a held-out ultra-sparse MoE with 12B total parameters and only 1/64 of its experts activated, the predicted hyperparameters remain close to the observed optima, supporting joint extrapolation across model scale and sparsity. Further experiments demonstrate transfer across expert granularities and isolate the effect of activation ratio from that of total expert count.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑