发表机构
University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨异构输入融合下的良性过拟合,发现其由联合信号与谱几何决定,而非仅由边际良性决定,且融合对回归和分类的影响可不同。
AI 中文摘要
良性过拟合在从单一高维输入学习时已被广泛研究,但其在异构输入融合下的行为在很大程度上仍未得到探索。我们在异构高斯设计下研究最小范数线性插值的这一问题,在保持底层总体任务不变的情况下,比较两个统计相关的输入块及其融合。对于回归,我们识别出一个全谱协方差证书,其渐近状态独立于截断阈值,并证明该证书被与两个边际分布一致的每个半正定联合协方差所保留。这种保护是严格的,但并未扩展到所有良性回归问题:在认证区域之外,两个良性边际分布可能产生有害的融合。对于单稀疏高斯分类,正则区域中的良性由存活的预测信号与噪声污染的平衡来刻画。融合可以使这两个量朝相反方向移动,并且在此模型类中,每个边际到联合的良性/非良性模式都是可实现的。我们进一步表明,相同的融合输入对回归和分类可以产生性质不同的影响。这些结果确立了异构融合下的良性过拟合由输入交互产生的联合信号和谱几何决定,而非仅由边际良性决定。
英文摘要
Benign overfitting is extensively studied when learning from a single high-dimensional input, but its behavior under heterogeneous input fusion remains largely unexplored. We study this question for minimum-norm linear interpolation under a heterogeneous Gaussian design, comparing two statistically dependent input blocks with their fusion while holding the underlying population task fixed. For regression, we identify a full-spectrum covariance certificate whose asymptotic status is independent of the cutoff threshold and prove that it is preserved by every positive-semidefinite joint covariance consistent with the two marginals. This protection is sharp, yet it does not extend to all benign regression problems: outside the certified regime, two benign marginals can have a harmful fusion. For one-sparse Gaussian classification, benignity in the regular regime is characterized by the balance between surviving predictive signal and nuisance contamination. Fusion can move these two quantities in opposite directions, and within this model class every marginal-to-joint benign/non-benign pattern is attainable. We further show that the same fused input can have qualitatively different effects on regression and classification. These results establish that benign overfitting under heterogeneous fusion is determined by the joint signal and spectral geometry created by input interaction, rather than by marginal benignity alone.
Comments51 pages, 6 figures