arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33047cs.LG

矩阵优化器中自适应牛顿-舒尔茨方法的零成本谱估计

Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers

Kristi Topollai, Anna Choromanska

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出零成本谱估计方法,利用牛顿-舒尔茨迭代中的格拉姆矩阵获取谱矩,实现谱自适应正交化,在固定迭代预算下降低误差或减少迭代次数,并提升GPT预训练中矩阵优化器的验证损失。

中文摘要 AI 辅助

诸如Muon之类的矩阵优化器通过近似正交化来变换每个动量矩阵,该近似正交化通常由少量牛顿-舒尔茨矩阵乘法实现。这种近似的质量和成本在很大程度上取决于其输入的奇异值谱,然而现有实现对于每一层以及整个训练过程都使用相同的固定多项式例程。我们表明这种统一处理是不必要的:牛顿-舒尔茨方法中的计算已经揭示了足够的信息,足以使该方法具有自适应性。牛顿-舒尔茨迭代内部形成的格拉姆矩阵通过廉价的标量归约产生谱矩,无需额外的矩阵乘法。从这些矩中,我们恢复经验奇异值分布的估计,并利用它选择专门针对当前矩阵的多项式例程。这将牛顿-舒尔茨正交化转变为一种谱自适应过程,能够响应不同层以及训练时间上的差异。在保存的动量矩阵上,谱估计在固定的迭代预算下显著降低了正交化误差,或者以更少的迭代达到相同的精度,并且在高达10亿参数的GPT预训练中,它降低了两种矩阵优化器的验证损失。我们的结果表明,优化器内部的矩阵函数运算无需为保守的最坏情况谱而设计:它们可以廉价地测量它们已经在处理的谱,并相应地专门化计算。

英文摘要

Matrix optimizers such as Muon transform each momentum matrix through an approximate orthogonalization, typically implemented by a small number of Newton-Schulz matrix multiplications. The quality and cost of this approximation depend strongly on the singular-value spectrum of its input, yet existing implementations use the same fixed polynomial routine for every layer and throughout training. We show that this uniform treatment is unnecessary: the computations in the Newton-Schulz method already reveal enough information to make the method adaptive. The Gram matrices formed inside Newton-Schulz iterations yield spectral moments through inexpensive scalar reductions, requiring no additional matrix multiplications. From these moments, we recover an estimate of the empirical singular-value distribution and use it to select a polynomial routine specialized to the current matrix. This turns Newton--Schulz orthogonalization into a spectrum-adaptive procedure that responds to differences across both layers and training time. On saved momentum matrices, spectral estimation substantially reduces orthogonalization error at a fixed iteration budget or reaches the same accuracy with fewer iterations, and in GPT pretraining up to 1B parameters it lowers the validation loss of two matrix optimizers. Our results suggest that matrix-function operations inside optimizers need not be designed for a conservative worst-case spectrum: they can cheaply measure the spectrum they are already processing and specialize computations accordingly.

发表机构

  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

↑