发表机构
National Key Laboratory for Novel Software Technology; School of Computer Science, Nanjing University(计算机软件新技术全国重点实验室; 南京大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出矩阵均衡Muon(MeqMuon)优化器,通过自动适应不平衡模式的归一化同时平衡行和列幅度,并省去AdamW二阶矩存储,在LLM预训练中实现更优收敛和更低内存占用。
AI 中文摘要
大语言模型(LLM)的成功伴随着模型规模和预训练成本的持续增长。Muon优化器在LLM预训练中展现出高精度和高训练效率。近期研究将逐行归一化引入Muon,以平衡更新幅度并提升预训练性能。然而,仅靠逐行归一化无法适应更新矩阵中不同的不平衡模式。本文提出一种改进的Muon优化器,称为矩阵均衡Muon(MeqMuon),用于LLM预训练。MeqMuon通过归一化同时平衡行和列的幅度,该归一化可自动适应不同的不平衡模式,无需人工干预。此外,MeqMuon无需存储AdamW的二阶矩估计,从而减少了优化器状态的内存占用。实验结果表明,在LLM预训练中,MeqMuon相比AdamW、Muon及其他基线取得了更好的收敛性能。
英文摘要
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.