发表机构
University of California, Los Angeles; The University of Hong Kong(加州大学洛杉矶分校; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过事实回忆模型揭示,Muon 的谱正交化将梯度下降的学习时间比从 Θ̃(√(S/R)) 降至 Θ̃(1),并加速误差指数衰减,同时保持正交等变性,从而更优地学习事实。
AI 中文摘要
Muon 优化器对矩阵值更新应用谱正交化,并在大规模神经网络训练中展现出强劲性能,然而这种变换在特征学习中的机制仍鲜为人知。在本工作中,我们通过一个可处理的事实回忆模型来研究这一问题,其中每个事实将主体-关系对映射到一个答案,而线性 Transformer 学习恢复该映射所需的主体依赖和关系依赖信息。该 Transformer 分别使用梯度流(GF)、谱梯度流(Spectral GF)或符号梯度流(Sign GF)进行优化,它们是梯度下降、Muon 和 Adam 的连续时间极限。先前研究(Nichani 等,2025)表明,当主体数量超过关系数量时,GF 先学习关系依赖信息,再学习主体依赖信息,从而在训练过程中产生特征分离阶段。我们用预测的主体依赖和关系依赖分量达到目标精度时的学习时间来刻画这一分离。对于 S 个主体和 R 个关系,GF 的学习时间比为 \widetilde{\Theta}(\sqrt{S/R}),而谱 GF 将该比降至 \widetilde{\Theta}(1)。此外,对于固定的 S 和 R,在 GF 下,主体依赖和关系依赖误差随训练时间 T 以 1/(T\log T) 衰减,而在谱 GF 下以 \exp(-\mathrm{poly}(T)) 衰减。最后,我们证明 GF 和谱 GF 在令牌嵌入的正交变换下是等变的,而符号 GF 则不然:不同的正交嵌入可能产生无特征分离、大的特征分离阶段,甚至颠倒的学习顺序。这些结果为谱正交化如何从根本上重塑特征学习动态提供了机制性视角。
英文摘要
The Muon optimizer applies spectral orthogonalization to matrix-valued updates and has shown strong performance in large-scale neural network training, yet the mechanisms of this transformation in feature learning remain poorly understood. In this work, we investigate this question through a tractable factual-recall model, where a fact maps each subject-relation pair to an answer, and a linear transformer learns the subject- and relation-dependent information required to recover this mapping. The transformer is optimized with gradient flow (GF), spectral GF, or Sign GF, which are continuous-time limits of gradient descent, Muon, and Adam, respectively. Prior studies (Nichani et al., 2025) have shown that when the number of subjects exceeds the number of relations, GF learns relation-dependent information before subject-dependent information, producing a feature-separation phase during training. We characterize this separation with the learning times when the subject- and relation-dependent components of the prediction reach a target accuracy. With $S$ subjects and $R$ relations, GF has a learning-time ratio of $\widetildeΘ(\sqrt{S/R})$, whereas Spectral GF reduces this ratio to $\widetildeΘ(1)$. In addition, for fixed $S$ and $R$, the subject- and relation-dependent errors decay as $1/(T\log T)$ in training time $T$ under GF, but as $\exp(-\mathrm{poly}(T))$ under spectral GF. Finally, we show that GF and spectral GF are equivariant under orthogonal transformations of the token embeddings, whereas Sign GF is not: Different orthonormal embeddings can potentially produce no feature separation, a large feature-separation phase, or even a reversed learning order. These results provide a mechanistic view of how spectral orthogonalization can fundamentally reshape feature-learning dynamics.
Comments47 pages, 8 figures