发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种适用于存在异方差和自相关的低维数据的岭回归估计量有限样本分布的高斯近似方法,并基于此提出两种正则化参数选择策略以优化预测风险。
AI 中文摘要
本文提出了一种针对经典岭回归估计量有限样本分布的简单高斯近似方法。该近似方法捕捉了有限样本下岭回归估计量通过权衡偏差与方差来降低估计和预测误差的特性,其基于非标准渐近理论:i) 令估计量的正则化参数与样本量成比例增长;ii) 将总体回归系数视为与定义估计量收缩方向的参考向量“局部”相关。与文献中其他渐近近似不同,该方法允许数据生成过程存在一般形式的异方差和自相关(代价是考虑低维模型,协变量数量不随样本量增长)。我们利用该简单高斯近似,提出两种选择岭回归估计量正则化参数的新策略,这些策略通过最小化平均或最坏情况的超额预测风险来选择正则化参数,其中风险采用本文提出的高斯近似计算。
英文摘要
We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator. Our approximation captures the fact that, in finite samples, the ridge regression estimator trades off bias and variance to reduce estimation and prediction error. Our approximation is based on nonstandard asymptotics where $i)$ we let the estimator's regularization parameter grow proportionally to the sample size; and $ii)$ we treat the population regression coefficients as \emph{local} to the reference vector that defines the estimator's direction of shrinkage. In contrast to other asymptotic approximations in the literature, we allow for general forms of heteroskedasticity and autocorrelation in the data generating process (at the cost of considering a low-dimensional model where the number of covariates is not allowed to grow with the sample size). We use our simple Gaussian approximation to propose two new strategies to select the regularization parameter for the ridge regression estimator. The suggested strategies select the regularization parameter to minimize either average or worst-case excess prediction risk, where risk is computed using our suggested Gaussian approximation.
Comments16 Figures