梯度噪声下的Muon与最优值附近正交化的极限
Muon Under Gradient Noise and the Limits of Orthogonalization Near Optima
AI总结:
本文研究Muon在噪声主导优化下的行为,发现其正交化引入非线性残差,增加协方差并提高损失下限,理论建议在噪声主导时用响应匹配的动量SGD替代正交化。
AI中文摘要:
Muon将每个权重矩阵的动量缓冲区替换为其正交极因子。我们研究这种正交化在最优值附近的作用,此时小批量噪声主导梯度。在高斯噪声模型下,预期的Muon更新变为缩放梯度步,因此具有适当匹配学习率的线性方法可以复现Muon的一阶平均响应。随机更新则不同:在响应匹配后,Muon保留了一个与输入噪声不相关的非线性残差,并贡献额外的协方差。Hermite展开展示了动量如何作用于该残差。其高阶分量比线性分量更快地去相关,因此动量相对于线性部分抑制了残差的累积协方差,但从未完全消除它。在局部二次代理中,残差在静态噪声缓冲区上评估,残差增加静态协方差并提高每个稳定步长下的静态损失下限,同时保持收缩动力学不变。在二次函数上的完整非线性递归模拟以及冻结Transformer梯度上的测量支持了这一图景的每一步。总之,这些结果为在优化变得噪声主导时,用响应匹配的动量SGD替代正交化提供了理论依据。
英文摘要:
Muon replaces the momentum buffer of each weight matrix by its orthogonal polar factor. We ask what this orthogonalization does near an optimum, where minibatch noise dominates the gradient. Under a Gaussian noise model, the expected Muon update becomes a scaled gradient step, so a linear method with a suitably matched learning rate reproduces Muon's first-order mean response. The stochastic update is a different matter: after the response is matched, Muon retains a nonlinear residual that is uncorrelated with the input noise and contributes additional covariance. A Hermite expansion shows how momentum acts on this residual. Its higher-order components decorrelate faster than the linear component, so momentum suppresses the residual's accumulated covariance relative to the linear part, but never eliminates it. In a local quadratic surrogate that evaluates the residual on the stationary noise buffer, the residual adds stationary covariance and raises the stationary loss floor at every stable step size, while leaving the contraction dynamics unchanged. Simulations of the full nonlinear recursion on quadratics and measurements on frozen transformer gradients support each step of this picture. Together, the results make a theoretical case for replacing orthogonalization by response-matched momentum SGD once optimization becomes noise-dominated.