惩罚性回归的嵌套交叉验证快速计算方法
Fast Computation of Nested Cross-Validation for Penalized Regression
AI总结:
本文提出一种仅需单次模型拟合的高效方法,可快速计算惩罚性回归(含岭回归等)的嵌套交叉验证预测区间,经函数型主成分回归实验验证,能大幅缩短运行时间。
AI中文摘要:
交叉验证是一种重抽样过程,可为任何预测模型提供泛化误差的点估计,广泛用于模型选择与评估。交叉验证估计的不确定性难以量化,其方差估计需多次运行完整重抽样过程。嵌套交叉验证通过对整个交叉验证过程重抽样,为给定模型和训练数据集计算泛化误差的预测区间,但计算成本极高。本文针对岭回归、样条平滑及部分函数型回归模型等惩罚性回归模型,提出一种仅需单次模型拟合即可计算嵌套交叉验证预测区间的高效方法。我们明确了在不同缩放 regime 及有限样本下,该方法优于基于重抽样的嵌套交叉验证的适用场景。针对函数型主成分回归的实验表明,在诸多重要场景中,该方法可大幅缩短运行时间。
英文摘要:
Cross-validation is a resampling procedure that provides a point estimate of generalization error for any predictive model. Cross-validation is widely used for model selection and evaluation. Uncertainty in the cross-validation estimate is challenging to quantify, and estimation of its variance is known to require multiple runs of the entire resampling procedure. Nested cross-validation computes a prediction interval for the generalization error for a given model and training dataset by resampling the entire cross-validation procedure but incurs extraordinary computational cost. We provide an efficient method for computing the nested cross-validation prediction interval using only a single model fit for some penalized regression models including ridge regression, spline smoothing, and some functional regression models. We characterize when our proposed method should be expected to out-perform resampling-based nested cross-validation in various scaling regimes as well as in finite samples. Experiments for functional principle components regression demonstrate non-trivial cases in which our proposed method improves run times substantially.