发表机构
Weizmann Institute of Science; University of Toronto; Vector Institute(魏茨曼科学研究所; 多伦多大学; Vector研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种分析一般范数与次高斯分布下线性回归良性过拟合的方法,证明p>1时最小p-范数插值可良性过拟合,而1-范数在非高斯分布下不成立,揭示现有结果对高斯性的依赖。
AI 中文摘要
理解预测器在插值噪声训练数据时仍能泛化的原因,是机器学习中的一个核心谜题。关于这种“良性过拟合”的大部分工作都研究最小2-范数线性回归,这反映了梯度下降的归纳偏置。然而,现代优化器如Adam和Muon使用非欧几里得更新几何,偏好与其他范数相关的解。分析非欧几里得范数下的回归要困难得多,已知结果基本上仅限于高斯分布。在本文中,我们开发了一种方法来分析一般范数和一般(次高斯)分布下线性回归中的良性过拟合。作为一个特例,我们证明了在适当条件下,对于p>1的最小p-范数插值,即使对于非高斯分布也能良性过拟合。也许令人惊讶的是,对于1-范数,良性过拟合在一般良态(但非高斯)分布下并不成立,这表明现有的1-范数正结果关键依赖于高斯性。我们的证明分析了其对偶优化问题的几何结构,使用集中和中心极限工具来表明在许多高维情况下该问题是近似欧几里得的。
英文摘要
Understanding why predictors can generalize despite interpolating noisy training data is a central puzzle in machine learning. Most work on such "benign overfitting" studies minimum-2-norm linear regression, reflecting the inductive bias of gradient descent. However, modern optimizers such as Adam and Muon use non-Euclidean update geometries, favoring solutions associated with other norms. Analyzing regression for non-Euclidean norms is substantially more difficult, with known results essentially limited to Gaussians. In this paper, we develop a method to analyze benign overfitting in linear regression for general norms and general (sub-Gaussian) distributions. As a special case, we prove that minimum-p-norm interpolation with p>1 can benignly overfit even for non-Gaussian distributions, under suitable conditions. Perhaps surprisingly, for the 1-norm, benign overfitting does not hold in general for well-behaved (but non-Gaussian) distributions, showing that existing positive 1-norm results rely crucially on Gaussianity. Our proof analyzes the geometry of the dual optimization problem, using concentration and central limit tools to show it is approximately Euclidean in many high-dimensional cases.