Muon 中的行归一化之谜
The Row Normalization Puzzle in Muon
浏览论文内容
中文总结 AI 辅助
本文研究 Muon 优化器中行归一化(NorMuon)的理论保证,发现其最坏情况复杂度存在维度因子,但实验显示其在 LLM 预训练中优于 Muon,揭示了理论与实践之间的谜团。
中文摘要 AI 辅助
本文研究了逐行重新归一化如何影响 Muon,重点关注 NorMuon 的最坏情况保证与其实际性能之间的差距(Li 等人)。尽管 NorMuon 在大语言模型(LLM)预训练中日益普及且性能表现良好,但其最坏情况保证仍鲜为人知。一个基本问题是:行归一化是否能带来可证明的收敛增益,可能通过其与近似极分解和指数移动平均动量的相互作用实现?我们的结果表明,在算子范数几何下,行归一化在最坏情况迭代复杂度中引入了一个依赖于维度的因子,即使采用精确极分解和任意固定动量参数,该因子仍然存在。实际上,我们在确定性设置中建立了算法相关的下界和匹配的上界,并将上界分析扩展到随机设置。两种上界分析均允许近似极分解。实验表明,在受我们最坏情况构造启发的合成问题上,NorMuon 比 Muon 慢,但在 LLM 预训练中却优于 Muon。这些发现加深了关于行归一化为何在实践中有效的谜团,并补充了 Dewulf 等人近期的发现。
英文摘要
This paper examines how row-wise renormalization affects Muon, focusing on the gap between NorMuon's worst-case guarantees and its practical performance (Li et al.). Despite its growing adoption and promising performance in large language model (LLM) pretraining, NorMuon's worst-case guarantees remain poorly understood. One fundamental question is: Does row normalization yield provable convergence gains, potentially through its interaction with approximate polar computation and exponential moving-average momentum? Our results show that row normalization introduces a dimension-dependent factor in the worst-case iteration complexity under the operator-norm geometry, which persists even with exact polar computation and any fixed momentum parameters. Indeed, we establish an algorithm-dependent lower bound and a matching upper bound in deterministic settings, and extend our upper bound analysis to stochastic settings. Both upper-bound analyses allow approximate polar computation. Experiments show that NorMuon is slower than Muon on synthetic problems inspired by our worst-case construction, yet outperforms Muon in LLM pretraining. These findings sharpen the puzzle of why row normalization helps in practice and complement the recent findings of Dewulf et al.
发表机构
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。