正交见证控制用于通过Sigmoid谱重塑的Muon优化
Orthogonal Witness Control for Muon Optimization via Sigmoid Spectral Reshaping
AI总结:
Soren是一种矩阵值优化器,通过sigmoid谱重塑保留梯度谱信息,在LLM预训练、微调和DPO中优于现有优化器。
AI中文摘要:
矩阵值优化器如Muon通过Newton-Schulz正交化利用神经网络更新的谱结构,但其对奇异谱的近扁平化处理丢弃了跨梯度模式的相对幅度信息。我们引入Soren(谱正交重塑),一种矩阵值优化器,它在保留梯度奇异子空间的同时,对其奇异值应用有界、单调的sigmoid变换。这平滑地压缩主导模式而不完全扁平化谱。我们将Soren解释为正定预条件梯度方法,并在相对光滑性和度量Polyak-Łojasiewicz几何下建立收敛保证。为避免显式奇异值分解,我们进一步开发了sigmoid谱映射的有限深度Soft Newton-Schulz(SNS)多项式实现,并表征其谱近似如何影响诱导的收敛几何。在LLM预训练、监督微调和直接偏好优化上的实验证明了Soren相对于既有优化器的有效性和鲁棒性。
英文摘要:
Matrix-valued optimizers such as Muon exploit the spectral structure of neural network updates through Newton--Schulz orthogonalization, but their near-flattening of the singular spectrum discards relative magnitude information across gradient modes. We introduce \emph{Soren} (\textbf{S}pectral \textbf{O}rthogonal \textbf{Re}shapi\textbf{n}g), a matrix-valued optimizer that preserves the singular subspaces of the gradient while applying a bounded, monotone sigmoid transformation to its singular values. This smoothly compresses dominant modes without fully flattening the spectrum. We interpret Soren as a positive-definite preconditioned gradient method and establish convergence guarantees under relative smoothness and metric Polyak--Łojasiewicz geometry. To avoid explicit singular value decomposition, we further develop a finite-depth Soft Newton--Schulz (SNS) polynomial realization of the sigmoid spectral map and characterize how its spectral approximation affects the induced convergence geometry. Experiments across LLM pre-training, supervised fine-tuning, and direct preference optimization demonstrate the effectiveness and robustness of Soren against established optimizers.