发表机构
University of California San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究训练损失平台期表征学习是否持续,发现矩阵Muon极化更新在损失平坦时仍使特征与教师子空间对齐,证明表征学习可与损失最小化解耦。
AI 中文摘要
当训练损失停止改善时,表征学习会停止吗?我们针对矩阵Muon研究了这一问题,其极化归一化更新的步长由梯度的秩而非其范数设定。在稳定性边缘附近,教师-学生问题上的全批量Muon进入近似周期为2的损失振荡,并持续数千步:周期平均损失保持平坦或上升,但权重持续移动,学习到的特征继续与教师子空间对齐。对于线性教师-学生学习玩具,我们推导了显式的周期和对齐公式以及条件性平台期和衰减界。对于群体平均场ReLU模型,我们证明在所述的维度、初始化和小型头部条件下,平均梯度外积(AGOP)的前导特征空间在损失平台期期间精确恢复教师子空间,随后损失下降。在我们研究的所有33种ReLU、GELU和SiLU教师配置中,仅方向对齐指标显示学生AGOP在周期2振荡期间与教师子空间对齐或仍在对齐;对选定配置的投影头部重拟合表明学习到的方向对预测有用,进一步的测量区分了AGOP对齐与权重质量集中。在深度残差ReLU学生中,冻结下游层而第一层使用全批量精确极化更新训练,重现了近乎平坦的周期平均损失,同时输入AGOP对齐改善;冻结和解冻在该平台期与损失下降之间切换,且该效应对方差和正交化器的选择敏感。
英文摘要
Does representation learning stop when the training loss stops improving? We study this question for matrix Muon, whose polar-normalised updates have a step length set by the gradient's rank rather than its norm. Near the edge of stability, full-batch Muon on teacher-student problems enters approximately period-2 loss oscillations that persist for thousands of steps: the cycle-mean loss stays flat or rises, yet the weights keep moving and the learned features continue to align with the teacher subspace. For linear teacher-student learning toys, we derive explicit cycle and alignment formulas and conditional plateau and decay bounds. For a population mean-field ReLU model, we prove that, under stated dimension, initialisation and small-head conditions, the leading eigenspace of the average gradient outer product (AGOP) recovers the teacher subspace exactly during a loss plateau, before the loss later drops. In all 33 ReLU, GELU and SiLU teacher configurations we study, direction-only alignment metrics show the student AGOP aligned with, or still aligning to, the teacher subspace during the period-2 oscillations; projected head refitting on selected configurations shows that the learned directions are useful for prediction, and further measurements distinguish AGOP alignment from weight-mass concentration. In deep residual ReLU students, freezing the downstream layers while the first layer trains with full-batch exact polar updates recreates a nearly flat cycle-mean loss with improving input-AGOP alignment; freezing and unfreezing switch between this plateau and loss decrease, and the effect is sensitive to momentum and to the choice of orthogonaliser.
Comments105 pages, 21 figures