AI 中文总结
研究高维线性插值器脆弱性,聚焦无脊回归与岭正则化估计器对比,用大偏差方法表明无脊回归风险有重尾行为,在大偏差率层面量化,发现其右尾衰减慢,揭示插值虽平均准确但可能统计脆弱,正则化影响风险事件频率。
AI 中文摘要
高维插值在现代机器学习中很常见,但其尾部风险不如预期预测风险那么为人所理解。现有理论表明插值模型预期表现良好,但这些保证并未确定罕见、严重错误的概率。在运筹学和随机决策应用中,罕见估计错误可能产生不成比例的下游影响,所以尾部行为与平均性能同样重要。我们使用大偏差方法研究高维线性插值器的脆弱性。聚焦无脊回归并将其与岭正则化估计器比较。首先表明无脊回归风险可能呈现重尾行为,尽管其预期风险可能得到很好控制,但其上尾衰减比正则化替代方案慢得多。然后在大偏差率层面量化此现象。在我们研究的情形中,岭正则化在\(n^2\)尺度抑制固定右尾偏差,而无脊回归只有\(n\log n\)尺度衰减,其中\(n\)是样本量。这种差距表明插值即使平均准确也可能在统计上很脆弱。因此正则化除了通常的偏差-方差权衡外,还影响罕见、高影响风险事件的频率。
英文摘要
High-dimensional interpolation is common in modern machine learning, but its tail risk is less understood than its expected prediction risk. Existing theory shows that interpolating models can perform well in expectation, yet such guarantees do not determine the probability of rare, severe errors. In operations research and stochastic decision-making applications, rare estimation errors can have disproportionate downstream effects, so tail behavior matters alongside average performance. We study the fragility of high-dimensional linear interpolators using large-deviation methods. We focus on ridgeless regression and compare it with ridge-regularized estimators. We first show that the risk of ridgeless regression can exhibit heavy-tailed behavior: although its expected risk may remain well controlled, its upper tail can decay much more slowly than that of regularized alternatives. We then quantify this phenomenon at the level of large-deviation rates. In the regime we study, ridge regularization suppresses fixed right-tail deviations at the $n^2$ scale, whereas ridgeless regression has only $n\log n$-scale decay, where $n$ is the sample size. This gap shows that interpolation can be statistically fragile even when it is accurate on average. Thus regularization affects the frequency of rare, high-impact risk events in addition to the usual bias-variance tradeoff.