arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17759math.OCcs.LG

无导数结构化更新用于Muon

Derivative-Free Structured Updates for Muon

  • Lawrence Berkeley National Laboratory(劳伦斯伯克利国家实验室)
  • University of California(加利福尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Pengcheng Xie

AI总结:

本文提出无导数框架,用结构化有限差分构造Muon风格更新,通过四种变体在矩阵回归和神经网络实验中验证,表明结构化探测在特定黑箱问题中实用,但无通用收敛保证。

AI中文摘要:

Muon通过正交化基于梯度的动量矩阵来更新矩阵值神经网络参数。其对导数的依赖限制了在梯度不可用或不可靠时的使用。我们开发了一个无导数框架,通过结构化有限差分构造Muon风格的更新。考虑了四种变体:全元素恢复、随机低秩代理、基对齐秩一探测和直接结构化搜索。穷举基对齐探测在理想极正交化之前的正缩放意义上等价于坐标有限差分。矩阵回归实验表明,随机秩一探测可以大幅减少函数评估次数,但代价是更新精度较低。在回归和神经网络上的受控噪声梯度实验说明了准确的函数值何时可以补偿不可靠的梯度预言机。一项小型CartPole研究进一步在固定回合预算下检验了正交秩一探测。这些结果支持结构化探测作为选定黑箱问题的实用选项;它们并未建立一般收敛保证或相对于准确、廉价梯度的优势。

英文摘要:

Muon updates matrix-valued neural-network parameters by orthogonalizing a gradient-based momentum matrix. Its reliance on derivatives limits its use when gradients are unavailable or unreliable. We develop a derivative-free framework that constructs Muon-style updates from structured finite differences. Four variants are considered: full entrywise recovery, random low-rank surrogates, basis-aligned rank-one probing, and direct structured search. Exhaustive basis-aligned probing is equivalent, up to positive scaling before ideal polar orthogonalization, to coordinate finite differences. Matrix-regression experiments show that random rank-one probing can reduce the number of function evaluations substantially, at the cost of less accurate updates. Controlled noisy-gradient experiments on regression and a neural network illustrate when accurate function values can compensate for an unreliable gradient oracle. A small CartPole study further examines orthogonal rank-one probes under a fixed episode budget. These results support structured probing as a practical option for selected black-box problems; they do not establish a general convergence guarantee or an advantage over accurate, inexpensive gradients.

补充信息

↑