arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MALT:通过对角预条件化实现的轻量曲率感知Muon

MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning

Tongle Wu, Huanyu Dong, Ying Sun, Ziye Ma

arXiv 2608.05088首次发表:更新:

发表机构

School of Electrical Engineering and Computer Science, The Pennsylvania State University; Department of Computer Science, City University of Hong Kong(宾夕法尼亚州立大学电气工程与计算机科学学院; 香港城市大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对Muon未考虑损失景观曲率几何的问题,提出轻量曲率感知优化方法MALT及MALTER,在GPT-2系列预训练上性能优于Muon且开销相近。

AI 中文摘要

Muon最近通过使用Newton-Schulz迭代对动量矩阵进行正交化,成为了AdamW在语言模型预训练中颇具前景的替代方案。尽管Muon缓解了梯度各向异性,但它并未明确考虑损失景观的曲率几何,因此可能仍对曲率各向异性敏感。为填补这一空白,我们提出MALT(Muon Augmented by Lightweight Two-sided Preconditioning,即通过轻量双边预条件化增强的Muon),它使用轻量对角预条件器降低Muon对曲率各向异性的敏感性。具体而言,MALT使用内存和计算开销低的双边对角预条件器,近似捕捉损失景观的曲率几何;它通过Newton-Schulz迭代对预条件化后的动量进行正交化,将结果映射回以定义更新方向,同时通过范数嫁接控制更新幅度。为提升MALT对随机梯度噪声的鲁棒性,我们进一步提出MALTER(MALT with Adaptive stEpsize Rescaling,即带自适应步长缩放的MALT)。我们为MALT在随机非凸场景下提供了收敛保证。在GPT-2 Small、Medium和Large预训练上的实验表明,所提方法在保持几乎相同内存占用和 wall-clock 时间的同时,性能优于Muon。

英文摘要

Muon has recently emerged as a promising alternative to AdamW for language model pretraining by orthogonalizing momentum matrices using Newton-Schulz iterations. Although Muon mitigates gradient anisotropy, it does not explicitly account for the curvature geometry of the loss landscape and may therefore remain sensitive to curvature anisotropy. We bridge this gap by proposing MALT (Muon Augmented by Lightweight Two-sided Preconditioning), which uses lightweight diagonal preconditioners to reduce the sensitivity of Muon to curvature anisotropy. Specifically, MALT uses two-sided diagonal preconditioners with low memory and computational overhead to approximately capture the curvature geometry of the loss landscape. It orthogonalizes the preconditioned momentum using Newton-Schulz iterations and maps the result back to define the update direction, while norm grafting controls the update magnitude. To improve the robustness of MALT to stochastic gradient noise, we further propose MALTER (MALT with Adaptive stEpsize Rescaling). Convergence guarantees are provided for MALT in the stochastic non-convex setting. Experiments on GPT-2 Small, Medium, and Large pretraining show that the proposed methods outperform Muon while maintaining nearly the same memory footprint and wall-clock time.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑