计算最优并非集群最优:面向稀疏混合专家模型的系统感知缩放
Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts
浏览论文内容
中文总结 AI 辅助
本研究开发MOSAIC框架,将稀疏MoE模型的架构与系统协同设计建模为优化问题,发现计算最优与集群最优稀疏性存在差异,主张统一架构与系统协同设计。
中文摘要 AI 辅助
在大规模预训练中,算法、架构与系统决策通常在不同阶段独立完成:缩放律阶段选择架构与训练方案,在计算约束下优化损失;单独的系统阶段则针对硬件效率优化实现。本研究中,我们开发了MOSAIC,它将模型架构与系统协同设计建模为一个优化问题。MOSAIC将预测缩放律与校准后的性能模型相结合,该模型可估计模型FLOPs利用率(MFU)、通信成本、内存占用及最优并行布局。我们将该框架实例化为稀疏混合专家(MoE)语言模型,其中专家数量、路由稀疏性及其他MoE层维度同时影响损失与系统效率。我们对在文本数据上训练的稀疏MoE模型拟合了缩放律,其缩放维度包括稀疏因子,即前向传播中每个token的非活跃模型参数占比。本研究中的缩放范围涵盖1.04亿个活跃参数至27亿个活跃参数,总模型规模达790亿个参数。我们表明,在校准后的稀疏性范围内,不考虑效率的模型FLOPs预算不存在内部最优稀疏性;拟合损失随模型稀疏度提升单调下降,计算最优位于数据支撑的上边界。而MoE模型的最优稀疏性是在集群的系统约束下产生的,这一点由MOSAIC捕捉。我们的结果支持在前沿语言模型训练中转向统一的架构与系统协同设计。
英文摘要
In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. A scaling law stage selects an architecture and training recipe, optimizing loss under compute constraints, and a separate systems stage then optimizes the implementation for hardware efficiency. In this work, we develop MOSAIC, which formulates model architecture and systems co-design as an optimization problem. MOSAIC couples a predictive scaling law with a calibrated performance model that estimates Model FLOPs Utilization (MFU), communication cost, memory footprint, and the best parallel layout. We instantiate the framework for sparse Mixture-of-Experts (MoE) language models, where expert count, routing sparsity, and other MoE layer dimensions affect both the loss and systems efficiency. We fit a scaling law on sparse MoE models trained on text data, whose scaling dimensions include the sparsity factor, which is the fraction of model parameters inactive per token in a forward pass. The scaling law sweeps in our work span active parameters from $104$ million to $2.7$ billion and total model sizes reaching $79$ billion parameters. We show that, within the calibrated sparsity range, an efficiency-agnostic model-FLOPs budget admits no interior optimal sparsity. The fitted loss decreases monotonically with sparser models and the compute optimum lies at the upper boundary of the data support. An optimal sparsity in MoE models instead emerges under the cluster's systems constraints, as captured by MOSAIC. Our results argue for a shift towards unified architecture and systems co-design for frontier language model training.
发表机构
- Amazon AGI Foundations(亚马逊AGI基础研究部门)
机构由 AI 辅助整理,请以论文原文为准。