arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20061cs.LGcs.AIcs.CL

逐步扩展:面向大规模混合专家模型的计算高效超参数迁移

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Nayeon Kim, Hojin Lee, Yunju Bak, Jaesun Park, Boseop Kim

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出两步超参数迁移框架,利用小型代理模型外推大规模混合专家模型的最优学习率,仅需极低成本即可准确预测全规模模型的最优配置。

中文摘要 AI 辅助

混合专家(MoE)架构可在不按比例增加计算成本的前提下显著扩展模型容量,然而在模型规模与 token 预算均达极端水平时,通过网格搜索优化其超参数(尤其是学习率)的计算成本极高。本文提出一种计算高效的两步超参数迁移框架,用于通过跨模型宽度迁移最优学习率,进而外推至万亿 token 训练场景,以估算大规模 MoE 模型训练的最优学习率。首先,我们针对采用多头潜注意力(MLA)与 Muon 优化器的 MoE 架构,构建最大更新参数化(μP)适配方案,证明最优学习率可在宽度缩放的模型间稳定迁移。其次,我们通过建立预测缩放定律,将该可迁移性扩展至 token 维度。通过对有限预算下小型代理模型的最优值进行线性回归,我们成功将理想学习率高保真度外推至大规模训练场景(如 10 万亿 token,R²=0.95)。这表明,对小型模型进行代理训练即可确定大规模 MoE 模型扩展训练的最优学习率。我们将所提方法应用于从头预训练基础模型(总参数 1550 亿,激活参数 170 亿),稳定的训练与评估结果验证,仅需极少的消融成本即可准确预测全规模目标模型的最优配置。

英文摘要

Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However, optimizing their hyperparameters---particularly the learning rate---at extreme scales of both model size and token budget via sweeping remains computationally prohibitive. In this paper, we propose a compute-efficient, two-step hyperparameter transfer framework that estimates optimal learning rates for training large MoE models by transferring them across scaling model widths, and subsequently extrapolating to trillion-token horizons. First, we formulate a Maximal Update Parameterization ($μ$P) adaptation for MoE architectures utilizing Multi-head Latent Attention (MLA) and the Muon optimizer, demonstrating that optimal learning rates transfer consistently across width-scaled models. Second, we extend this transferability along the token dimension by establishing a predictive scaling law. By applying linear regression to the optimal values derived from small proxy models on limited budgets, we successfully extrapolate the ideal learning rate to massive training horizons (e.g., 10 trillion tokens) with high fidelity ($R^2=0.95$). Consequently, this indicates that proxy training on small models is sufficient to determine the optimal learning rate for the extensive training of large-scale MoEs. We apply the proposed methodology to pretrain our foundation model (155B total, 17B active parameters) from scratch, and the stable training and evaluation results validate that optimal configurations for full-scale target models can be accurately predicted with minimal ablation costs.

发表机构

  • Kakao Corp.(Kakao公司)
  • Upstage AI(Upstage人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑