最大检验校准的Stein收缩与超高维回归中的诚实子模型选择
Max-Test-Calibrated Stein Shrinkage with Honest Submodel Selection in Ultra-High-Dimensional Regression
浏览论文内容
中文总结 AI 辅助
本文提出一个样本分离的诚实框架,利用最大残差关联统计量和投影高斯抽样校准Stein型收缩估计量,在超高维回归中实现有效的子模型选择与风险控制。
中文摘要 AI 辅助
经典的事前检验和Stein型估计量在受限回归与全回归拟合之间进行插值,但当$p\gg n$且限制是数据自适应时,OLS和卡方校准失效。我们提出一个诚实的样本分离框架,其核心是一个共同的选定零假设分布。独立的选定数据扩展了一个强制性核心;独立的估计数据将完整设计的正则化全模型(FM)拟合与全新的精确零假设子模型重拟合配对。一个最大残差关联统计量评估所有被排除的坐标。投影高斯抽样在同方差高斯误差下再现其条件零分布,产生一个有限样本有效的秩检验。同一设计特定分布的逆矩取代$q-2$并校准事前检验(PT)、Stein型(S)和正部分Stein型(PS)规则。我们推导出精确端点风险公式、诚实条件有效性、零分布集中性、逆矩一致性、显式选择失败余项以及端点自适应性,而不声称均匀支配性。高斯实验进行了2000次重复,预测变量多达30000个,比较了Ridge和LASSO FM,通过冻结折交叉验证调参,在共同的子模型、检验和校准下进行。PS在接近零的系数下提供较大的风险增益,并在强偏离下接近相关的全模型风险。一个独立的CPSS-LASSO/CPSS-MCP审计评估数据自适应选择,而一个包含19152个预测变量的分裂样本DepMap研究说明了预测增益和对误差假设的敏感性。HDMaxShrink R包实现了该程序。
英文摘要
Classical preliminary-test and Stein-type estimators interpolate between restricted and full regression fits, but OLS and chi-squared calibration fail when $p\gg n$ and the restriction is data-adaptive. We propose an honest sample-separated framework built around a common selected-null law. Independent selection data extend a mandatory core; separate estimation data pair a complete-design regularized full-model (FM) fit with a fresh exact-null submodel refit. A maximum residual-association statistic assesses all excluded coordinates. Projected Gaussian draws reproduce its conditional null distribution under homoskedastic Gaussian errors, yielding a finite-sample-valid rank test. The inverse moment of the same design-specific law replaces $q-2$ and calibrates preliminary-test (PT), Stein-type (S), and positive-part Stein-type (PS) rules. We derive exact endpoint-risk formulas, honest conditional validity, null-law concentration, inverse-moment consistency, an explicit selection-failure remainder, and endpoint adaptivity without uniform-dominance claims. Gaussian experiments with 2,000 replications and up to 30,000 predictors compare Ridge and LASSO FMs, tuned by frozen-fold cross-validation, under a common submodel, test, and calibration. PS provides large near-null coefficient-risk gains and approaches the relevant full-model risk under strong departures. A separate CPSS-LASSO/CPSS-MCP audit evaluates data-adaptive selection, while a split-sample DepMap study with 19,152 predictors illustrates prediction gains and sensitivity to error assumptions. The HDMaxShrink R package implements the procedure.