发表机构
University of California-San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明神经网络在种群损失平台期仍能学习更优表示,通过理论分析和实验展示了教师子空间与AGOP特征空间对齐的显著提升及重拟合误差的降低。
AI 中文摘要
当神经网络学习到显著更具预测性的表示时,种群损失可以保持几乎恒定。我们针对在高斯输入上通过所有参数的同时固定步长种群梯度下降训练的两层ReLU和leaky-ReLU网络建立了这种分离。对于结构化加性教师,其链接是$H^1(\gamma)$中高斯阻尼三次函数的正混合,我们给出了显式条件,在这些条件下,小的IID高斯初始化产生高概率保证:在高损失平台期的一个检查点,秩-$r$教师子空间与预测器平均梯度外积(AGOP)的前$r$维特征空间之间的最小对齐相对于初始化增加至少$1/2$,并且在未改变的系数预算下的最小重拟合MSE相对于初始化减少超过$0.399$。相同的轨迹随后达到低于平台窗口中每个值的训练损失。一个补充结果使用投影特征重拟合处理不等权重三次教师和小加性Sobolev扰动。对于具有精确拟合截距的SwiGLU网络,我们证明在高斯初始化消失时,在固定宽度和维度下,损失平台期内的前导AGOP对齐,适用于具有非零一、二或三次Hermite内容的平方可积教师。一个秩一三次特化也在规定宽度下给出同时无限制重拟合增益。一个近似下界进一步表明,当脊神经元被限制在教师子空间内的共享正交轴上时,某些交互目标保留非零误差。使用ReLU学生在21个教师和每个教师50次初始化上的种群矩实验补充了分析。
英文摘要
Population loss can remain nearly constant while a neural network learns a substantially more predictive representation. We establish this separation for two-layer ReLU and leaky-ReLU networks trained on Gaussian inputs by simultaneous fixed-step population gradient descent on all parameters. For structured additive teachers whose links are positive mixtures of Gaussian-damped cubics in $H^1(γ)$, we give explicit conditions under which small IID Gaussian initialization yields a high-probability guarantee: at a checkpoint during a high-loss plateau, minimum alignment between the rank-$r$ teacher subspace and the leading $r$-dimensional eigenspace of the predictor's average gradient outer product (AGOP) increases by at least $1/2$, and the minimum refit MSE under unchanged coefficient budgets decreases by more than $0.399$, both relative to initialization. The same trajectory subsequently attains a trained loss below every value in the plateau window. A complementary result treats unequal-weight cubic teachers and small additive Sobolev perturbations using projected-feature refits. For SwiGLU networks with an exactly fitted intercept, we prove leading-AGOP alignment during a loss plateau at fixed width and dimension as Gaussian initialization vanishes, for square-integrable teachers with nonzero Hermite content of degree one, two, or three. A rank-one cubic specialization also gives simultaneous unrestricted-refit gains at a prescribed width. An approximation lower bound further shows that certain interaction targets retain nonzero error when ridge neurons are restricted to shared orthogonal axes within the teacher subspace. Population-moment experiments with ReLU students across 21 teachers and 50 initializations per teacher complement the analysis.
Comments127 pages, 27 figures