arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谱收缩用于Muon与半定规划中免分解矩阵滤波

Spectral Deflation for Factorization-Free Matrix Filtering in Muon and Semidefinite Programming

Haoran Sun, Shucheng Kang, Heng Yang

arXiv 2609.21102首次发表:更新:

发表机构

Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对Muon优化器和半定规划中滤波器在归一化后精度下降的问题,提出谱收缩方法,通过移除主导谱分量再滤波恢复,严格减小Newton-Schulz误差,并在GPT-2预训练和大规模SDP上提升性能。

AI 中文摘要

Muon优化器与一阶半定规划(SDP)求解器的GPU实现共享一种计算模式:矩阵分解被替换为归一化后施加的固定深度多项式滤波器。当少数主导谱分量承载了输入的大部分尺度时,归一化将剩余谱推向零,而滤波器在该区域精度最低。我们提出谱收缩:估计主导分量,移除它们,用未改变的滤波器对归一化残差进行滤波,然后恢复这些分量。收缩保留了目标矩阵函数,并且我们证明它严格减小了经典Newton-Schulz族的有限步误差。通过为Muon实现的批量随机SVD和为ADMM实现的热启动子空间跟踪,收缩在GPT-2预训练中持续优于相应的Muon基线(使用Newton-Schulz和Polar Express映射),在墙钟时间上也是如此,并在类似的投影时间内降低了大规模SDP上免分解ADMM的KKT残差。

英文摘要

GPU implementations of the Muon optimizer and of first-order semidefinite programming (SDP) solvers share one computational pattern: a matrix factorization is replaced by a fixed-depth polynomial filter applied after normalization. When a few dominant spectral components carry most of the input's scale, normalization pushes the remaining spectrum toward zero, where the filter is least accurate. We propose spectral deflation: estimate the dominant components, remove them, filter the normalized residual with the unchanged filter, and restore them. Deflation preserves the target matrix function, and we prove that it strictly reduces the finite-step error of the classical Newton--Schulz family. Implemented with batched randomized SVD for Muon and warm-started subspace tracking for ADMM, deflation consistently improves GPT-2 pretraining over the corresponding Muon baselines with both the Newton--Schulz and Polar Express mappings, also in wall-clock time, and lowers the KKT residuals of factorization-free ADMM on large-scale SDPs within a similar projection time.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑