arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13677math.OCcs.LG

非光滑优化中的正交化动量

Nonsmooth Optimization via Orthogonalized Momentum

Lexiao Lai, Tianyi Lin, Jiayu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究正交化动量在非光滑优化中的有效性,发现固定动量因子下Muon可能失效,提出自适应动量因子恢复收敛,并设计MAGD算法,在凸与非凸设置下均保证收敛,实验验证其作为Muon的实用替代方案。

中文摘要 AI 辅助

现代实际应用问题涉及矩阵值参数,然而传统优化器将其视为向量,这促使了利用输入输出几何结构的矩阵感知方法的发展,例如Muon,它在参数更新前对动量矩阵进行正交化。其经验上的成功引发了一个概念性问题:正交化动量在光滑优化之外是否仍然有效?本文在局部Lipschitz函数上,使用与反向传播兼容的广义导数框架来研究此问题。我们的第一个贡献是识别出一个关键限制:对于每个固定的动量因子$\beta\in[0,1)$,当步长适应全梯度历史时,Muon从几乎所有的初始化出发,都可能无法逼近凸Lipschitz目标的全局最优解。这种失败甚至可能发生在有界迭代序列中。我们的例子受到Parshakova等人的启发,但他们的结果仅覆盖$\beta\in[0,\frac{1}{2})$的情况。接着,我们表明障碍在于固定动量而非正交化。实际上,当动量因子自适应并随着步长趋于零而接近1时,在有界性和正则性条件下,Muon恢复了非凸非光滑优化的渐近收敛性。此外,我们提出了MAGD,它结合了正交化动量与梯度,并根据它们的相对进展进行加权。MAGD在非凸设置中保持渐近收敛性,并在凸设置中实现了$O(\min\{m,n\}\epsilon^{-2})$的速率。一个下界表明了最优的维度依赖性。在合成问题、图像分类和大语言模型预训练上的实验表明,MAGD是Muon的一个简单且实用的替代方案。综合起来,我们的结果刻画了正交化动量在缺乏光滑性时何时失效,以及如何使其变得可靠,我们希望这一分析能更广泛地发挥作用。

英文摘要

Modern real application problems involve matrix-valued parameters, yet conventional optimizers treat them as vectors, thereby motivating matrix-aware methods that exploit input-output geometry, such as Muon which orthogonalizes the momentum matrices before parameter updates. Its empirical success raises a conceptual question: can orthogonalized momentum remain effective beyond smooth optimization? This paper studies this question for locally Lipschitz functions using a generalized derivative framework compatible with backpropagation. Our first contribution is to identify a key limitation: for every fixed momentum factor $β\in[0,1)$, Muon can fail to approach the global optimal solution of a convex Lipschitz objective from almost every initialization, when step sizes adapt to the full gradient history. The failure can occur even along bounded iterates. Our example is inspired by the one of Parshakova et al. which only covers $β\in[0,\frac{1}{2})$. Then, we show that the obstruction lies in fixed momentum rather than orthogonalization. Indeed, when the momentum factor is adaptive and approaches 1 together with a vanishing step size, Muon recovers asymptotic convergence for nonconvex nonsmooth optimization under the boundedness and regularity conditions. Moreover, we propose MAGD, which combines orthogonalized momentum with gradient, weighted based on their relative progress. MAGD retains asymptotic convergence in nonconvex settings and achieves an $O(\min\{m,n\}ε^{-2})$ rate in convex settings. A lower bound shows the optimal dimension dependence. Experiments on synthetic problems, image classification, and LLM pretraining show MAGD is a simple and practical alternative to Muon. Together, our results characterize when orthogonalized momentum fails without smoothness and how it can be made reliable and we hope that the analysis may be useful more broadly.

发表机构

  • Columbia University(哥伦比亚大学)
  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑