AI 中文总结
本文针对Cox-Lasso变量选择后的推断难题,提出基于模型的自助推断方法,证明其一阶有效性,模拟和实例显示该方法能改善小样本条件覆盖率,为生存分析提供可解释的不确定性量化。
AI 中文摘要
Cox回归中变量选择后的推断颇具挑战,因为选择后简单的Wald型区间在有限样本下的条件覆盖率可能较差。本文研究了Cox-Lasso变量选择后的基于模型的自助推断方法:首先对原始数据拟合一次Cox-Lasso以选择变量集,随后仅使用这些变量拟合无惩罚的Cox模型;自助样本由半参数插件Cox模型生成,该模型由此无惩罚Cox重拟合的系数估计、Breslow基线累积风险估计量以及插件删失分布指定。在每个自助样本中,所选变量集保持固定,仅重拟合无惩罚Cox模型。在 oracle 型稀疏模型假设和标准Cox模型正则条件下,本文证明了该方法的一阶自助有效性。在考虑的模拟场景中,百分位自助区间和学生化自助区间在多个小样本和中等样本设置下,相比自助-Wald区间表现出更优的条件覆盖率;其性能与去偏区间大致相当,不过比较结果取决于信号强度、调参及选择稳定性。SEER乳腺癌实例表明,该方法可应用于实际生存分析,并为变量选择后报告的效应提供可解释的不确定性量化。
英文摘要
Inference after variable selection in Cox regression is difficult because simple Wald-type intervals after selection can have poor finite-sample conditional coverage. We study a model-based bootstrap for inference after Cox-Lasso variable selection. The Cox-Lasso is fitted once to the original data to select a set of variables, after which an unpenalized Cox model is fitted using only those variables. Bootstrap samples are generated from a semiparametric plug-in Cox model specified by the coefficient estimate from this unpenalized Cox refit, the Breslow baseline cumulative hazard estimator, and a plug-in censoring distribution. In every bootstrap sample, the selected variable set is kept fixed and only the unpenalized Cox model is refitted. Under oracle-type sparse-model assumptions and standard Cox model regularity conditions, we prove first-order bootstrap validity for this procedure. In the simulation scenarios considered, percentile and studentized bootstrap intervals showed improved conditional coverage relative to the bootstrap-Wald interval in several small- and moderate-sample settings. Their performance was broadly competitive with debiased intervals, although the comparison depended on signal strength, tuning, and selection stability. A SEER breast cancer example illustrates that the procedure can be implemented in a realistic survival analysis and provides interpretable uncertainty quantification for effects reported after variable selection.