arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新评估用于矩阵分解的Muon优化器

Reassessing Muon for Matrix Factorization

Ali Parviz, Gal Mishne, Alex Cloninger

arXiv 2607.13246首次发表:更新:

发表机构

Halicioğlu Data Science Institute, UC San Diego; Department of Mathematics, UC San Diego(加州大学圣地亚哥分校哈利西奥卢数据科学研究所; 加州大学圣地亚哥分校数学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在低秩矩阵分解问题上重新评估Muon优化器,通过与精心调整的自适应基线对比,发现Muon在此设置下不始终优于AdamW,一些优势对超参数选择敏感,更细致地展现谱感知正交化何时有益,主张在受控问题上评估优化器。

AI 中文摘要

Muon最近成为大规模深度学习的强大优化器,通过近似正交化重塑梯度更新,在大语言模型训练中表现优于Adam和AdamW。其经验成功激发了大量理论工作,将其解释为谱范数下的最速下降。但尚不清楚Muon的哪些优势源于其更新规则本身,哪些是现代深度网络规模、架构和数据的产物。本文通过在简单、易于理解且具有谱结构的低秩矩阵分解问题上研究Muon,将优化器与这些混杂因素分离。通过与精心调整的自适应基线进行对比,发现Muon在此设置下并不始终优于AdamW,且一些先前报告的优势对超参数选择敏感。结果更细致地描绘了谱感知正交化何时有益,并主张除了端到端基准测试外,还应在受控问题上评估现代优化器。

英文摘要

Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretical work that interprets Muon as steepest descent under the spectral norm. Yet it remains unclear which of Muon's advantages stem from its update rule itself and which are artifacts of the scale, architecture, and data of modern deep networks. In this work, we isolate the optimizer from these confounding factors by studying Muon on a simple, well-understood, and spectrally structured problem: low-rank matrix factorization. Through a controlled comparison against carefully tuned adaptive baselines, we find that Muon does not consistently outperform AdamW in this setting and that several previously reported advantages are sensitive to hyperparameter choices. Our results provide a more nuanced picture of when spectrum-aware orthogonalization is beneficial and argue for evaluating modern optimizers on controlled problems in addition to end-to-end benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑