发表机构
RIKEN AIP; University of Toronto; University of Tokyo; Kyoto University(理化学研究所先进智能项目中心; 多伦多大学; 东京大学; 京都大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究谱梯度下降在过参数化矩阵分类中的泛化,发现捷径几何(聚集或分散)可逆转其与GD的相对优势,并揭示单步SpecGD即可良好泛化而持续训练反而恶化。
AI 中文摘要
我们研究了在带有损坏标签的过参数化矩阵分类中,谱梯度下降(SpecGD)的泛化性能。每个输入由一个共享的低秩信号和一个秩一的样本特定扰动(称为捷径)组合而成,该扰动能够实现记忆但无法泛化。我们对比了共享单一奇异方向的“聚集捷径”与占据不同奇异方向的“分散捷径”。仅改变这种几何结构就可能逆转GD和SpecGD的相对泛化性能:聚集捷径可能有利于SpecGD,而分散捷径可能有利于GD。在分散机制中,捷径的完全正交性消除了SpecGD后期方向中的信号,而消失的随机相关性通过二阶效应共同产生一个微小但泛化相关的信号。为了识别SpecGD选择的方向(谱最大间隔问题本身无法确定),我们将其对偶的精细分析与归一化损失权重的指数梯度动力学相结合。最后,我们表明,单步SpecGD已经能够插值并良好泛化,而持续训练会收敛到一个泛化性能显著更差的方向。
英文摘要
We study the generalization of spectral gradient descent (SpecGD) in overparameterized matrix classification with corrupted labels. Each input combines a shared low-rank signal with a rank-one sample-specific perturbation, referred to as a shortcut, that enables memorization but does not generalize. We contrast collapsed shortcuts, which share a singular direction, with dispersed shortcuts, which occupy distinct singular directions. Changing only this geometry can reverse the relative generalization of GD and SpecGD: collapsed shortcuts can favor SpecGD, while dispersed shortcuts can favor GD. In the dispersed regime, exact shortcut orthogonality eliminates the signal from the late-stage SpecGD direction, while vanishing random correlations collectively generate a small but generalization-relevant signal through a second-order effect. To identify the direction selected by SpecGD, which the spectral max-margin problem alone does not determine, we combine a refined analysis of its dual with the exponentiated-gradient dynamics of normalized loss weights. Finally, we show that a single SpecGD step can already interpolate and generalize well, while continued training converges to a direction with substantially worse generalization.