AI 中文总结
本研究在(d+1)维单尖峰线性回归模型中发现“良性失配(第四象限)”区间,此时泛化好的预测器训练拟合差于零预测器,大学习率单轮SGD可在该区间达到接近最优的测试误差,且干扰分量同时控制训练失配与对抗敏感性。
AI 中文摘要
训练误差是我们在训练集上可以观测到的量;测试误差才是我们真正关心的指标。我们在确定性的(d+1)维单尖峰模型中研究带平方误差的线性回归。每个程式化训练向量都具有相同的信息性尖峰坐标,其幅度为√γ且γ>1。其余方向为干扰方向,不同训练向量的干扰分量范数相等且两两正交。训练标签均为1。新的测试点服从x_test ~ N(0, diag(γ,1,…,1))分布,无噪声测试标签为归一化尖峰坐标x_test[1]/√γ。我们聚焦于训练向量张成空间内的线性预测器——这是零初始化线性梯度方法自然得到的假设类。\n我们展示了存在一个训练集规模n的区间,其中所有泛化性能良好的张成空间预测器对训练数据的拟合效果都必然比零预测器更差。我们将这种机制称为“良性失配”,也叫第四象限。当n≫d/γ²时,最优张成空间预测器开始具备泛化能力,而插值法直到更晚的阈值n≫d/γ时才具备泛化能力。在d/γ² ≪ n ≪ d/γ的区间内,线性张成空间内的有效预测超出了插值范畴:训练点上的预测值会超过标签。我们证明,采用大恒定学习率的单轮随机梯度下降(SGD)在整个区间内都能达到较低的测试误差——与最优张成空间预测器的性能仅相差一个对数因子。我们还直接验证了该方法确实具有较大的经验训练误差(尽管其名称中带有“下降”的设定)。最后,我们证明导致训练失配的不可避免的干扰分量,同时也控制着预测器的对抗敏感性。
英文摘要
Training error is what we can observe on a training set; test error is the quantity we actually care about. We study linear regression with squared-error in a deterministic $(d+1)$-dimensional single-spike model. Each stylized training vector has the same informative spike coordinate, of amplitude $\sqrtγ$ with $γ>1$. The remaining directions are nuisance, and the nuisance components of distinct training vectors all have equal norm and are mutually orthogonal. The training labels are all $1$. Fresh test points are drawn from $\vec{x}_{\rm test} \sim \mathcal{N}(\vec{0},\operatorname{diag}(γ,1,\ldots,1))$, with the noise-free test labels being the normalized spike coordinate $x_{\rm test}[1]/\sqrtγ$. We focus on linear predictors in the span of the training vectors, the class naturally reached by zero-initialized linear gradient methods. We exhibit a range of training-set sizes $n$ in which every span predictor that generalizes well must fit the training data \emph{worse} than the zero predictor. We call this regime \emph{benign misfitting}, or the fourth quadrant. The best span predictor begins to generalize when $n\gg d/γ^2$, while interpolation does not generalize until the later threshold $n\gg d/γ$. In the window $d/γ^2 \ll n \ll d/γ$, useful prediction within the linear span lies beyond interpolation: predictions on the training points overshoot the labels. We show that one-pass stochastic gradient descent (SGD), with a large constant learning rate, reaches small test error throughout this window---matching the best span predictor up to a logarithmic factor. We also verify directly that it indeed has \emph{large} empirical training error (despite the descent premise in its name). Finally, we show that the unavoidable nuisance component responsible for the training misfit also controls the predictor's adversarial sensitivity.
Comments82 pages, 6 figures, full version of paper accepted at ITW2026