Muon 在 LLM 预训练中超越边缘稳定性
Muon Sublates the Edge of Stability in LLM Pretraining
浏览论文内容
中文总结 AI 辅助
Muon 优化器在 LLM 预训练中打破经典边缘稳定性耦合,实验表明其损失中性边界与时间对齐独立,支持分裂 EoS 图景。
中文摘要 AI 辅助
Muon 越来越多地被用于语言模型预训练,但其大步长动力学并未被经典的梯度下降(GD)边缘稳定性(EoS)图景所涵盖。在 GD 中,损失中性、等幅更新反转和边际稳定性在单一依赖于学习率的边缘处汇合。我们证明 Muon 打破了这种耦合。对于随机无动量 Muon,我们推导出一个相干校正的条件损失中性边界 $2\rho_b/\eta$,而时间对齐遵循独立的几何结构。受控实验表明,损失平衡和时间对齐对学习率和批大小的响应不同。在我们的语言模型实验中,130M Llama 类 LLM 运行表现出损失边界跟踪和弱负对齐,而所研究的 1B LLM 配置显示出更强的部分抵消;在两种设置中,方向仍远非相干反转,而训练持续改进。这些结果支持 Muon 的分裂 EoS 图景:随机损失中性边缘仍然存在,但并未伴随普遍的时间方向特征。用于复现实验的源代码可在此 https URL 中找到。
英文摘要
Muon is increasingly used for language-model pretraining, yet its large-step dynamics are not captured by the classical edge-of-stability (EoS) picture of gradient descent (GD). In GD, loss neutrality, equal-magnitude update reversal, and marginal stability meet at a single learning-rate-dependent edge. We show that Muon breaks this coupling. For stochastic no-momentum Muon, we derive a coherence-corrected conditional loss-neutral boundary $2ρ_b/η$, while temporal alignment follows a separate geometry. Controlled experiments show that loss balance and temporal alignment respond differently to learning rate and batch size. Across our language model experiments, the 130M Llama-like LLM runs exhibit loss-boundary tracking with weak negative alignment, whereas the studied 1B LLM configuration shows stronger partial cancellation; in both settings, directions remain far from coherent reversal while training continues to improve. These results support a split EoS picture for Muon: a stochastic loss-neutral edge survives, but it is not accompanied by a universal temporal-direction signature. The source code for reproducing the experiments can be found in https://github.com/cyzebra/Muon-Sublates-the-Edge-of-Stability-in-LLM-Pretraining
发表机构
- School of Mathematical Sciences, Shanghai Jiao Tong University(上海交通大学数学科学学院)
- Alibaba Inc.(阿里巴巴集团)
- School of Mathematical Sciences, Institute of Natural Sciences and MOE-LSC, Shanghai Jiao Tong University(上海交通大学数学科学学院、自然科学研究院及教育部重点实验室)
机构由 AI 辅助整理,请以论文原文为准。