AI 中文总结
研究针对选择赢家后估计受赢家诅咒影响的问题,引入灵活条件推断方法,通过自适应指数随机化校正过度乐观,实现与标准前 k 选择相近的质量且置信区间更短,广泛适用于多种非参数设置。
AI 中文摘要
研究人员常基于数据驱动标准选择表现最佳的选项或赢家,如处理方式、模型或模型特征,然后报告所选赢家的效应估计。然而,朴素的选择后估计会受赢家诅咒影响,产生系统性过度乐观的结果。我们引入一种灵活的条件推断方法,通过自适应指数随机化方案校正这种过度乐观。我们的方法实现的选择质量与标准的前 k 选择相近,同时置信区间比现有方法更短。此外,我们的方法广泛适用于具有渐近线性选择统计量的非参数设置,涵盖临床试验中最有前景治疗方法的疗效推断、排行榜上顶级模型的能力以及模型中最具预测性特征的重要性等广泛应用。
英文摘要
Researchers often select top-performing options or winners, based on a data-driven criterion, such as treatments, models, or model features and then report effect estimates for the selected winners. Naive post-selection estimates, however, are known to suffer from the winner's curse, producing systematically overoptimistic results. We introduce a flexible conditional inference method that corrects for this overoptimism through an adaptive exponential randomization scheme. Our method achieves selection quality that closely matches that of standard top-k selection, while also yielding shorter confidence intervals than existing approaches. Furthermore, our approach applies broadly to nonparametric settings with asymptotically linear selection statistics, covering wide-ranging applications such as inference for the efficacy of the most promising treatments in clinical trials, the abilities of top-ranked models on leaderboards, and the importance of the most predictive features in a model.