替代指标何时有效?现实失效模式下的有限样本比较
When Do Surrogate Metrics Work? A Finite-Sample Comparison Under Realistic Failure Modes
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过有限样本比较七种估计器,发现替代指标在替代性成立时高效但违反时覆盖率崩溃,而PPI++在随机标记下更稳健,建议将其作为主要分析并辅以诊断工具。
AI中文摘要:
在线实验通常必须在长期结果成熟之前进行评估。在滚动入组的情况下,这些结果仅对早期入组者可见,而短期替代指标对所有参与者均可用。我们比较了七种估计器,涵盖十一种数据生成过程,包括部分中介、漂移、结果稀疏性和入组时间标记,在超过500个方法-场景组合上进行了多达R=2,000次重复实验。我们发现了一个明显的稳健性-效率权衡:当替代性成立时,替代指标带来较大的效率提升,但在违反替代性时其覆盖率崩溃,而PPI族方法在随机标记下保持渐近有效性,但增益较小。我们给出了固定预测器下全单元参数化中PPI++的有限样本方差、交叉拟合下两种估计器的联合渐近分布,以及一个Hausman型估计器不一致诊断,然后量化了检测-损害差距:在部分中介设计中,我们定位到两个边界,一条违反带破坏了替代指标的覆盖率,但在大多数数据集上却太小而无法检测。在64,000名客户的Hillstrom实验中,诊断很少标记出会使替代指标产生偏差的违反,而两种估计器的Cauchy核混合继承了10.7%的相对偏差;在1400万用户的Criteo实验中,违反被检测到,子采样追踪检测随规模增加而开启,同时损害持续存在。PPI++存在局限性:在稀有转化n=30,000的Criteo子样本中,其经验覆盖率为85.0%。我们建议在随机标记和足够的标记结果计数下,将具有精确方差的PPI++预先指定为主要分析,将诊断视为警告而非证书,并将替代指标视为敏感性分析。
英文摘要:
Online experiments must often be evaluated before long-term outcomes mature. Under rolling enrollment, these outcomes are observed only for early enrollees, while short-term surrogates are available for everyone. We compare seven estimators across eleven data-generating processes, spanning partial mediation, drift, outcome sparsity, and enrollment-time labeling, with up to $R = 2,000$ replications over more than 500 method-by-scenario cells. We find a sharp robustness-efficiency tradeoff: the surrogate index delivers large efficiency gains when surrogacy holds but its coverage collapses under violations, while PPI-family methods stay asymptotically valid under random labeling at smaller gains. We give the finite-sample variance of PPI++ in the all-units parameterization for a fixed predictor, a joint asymptotic distribution for the two estimators under cross-fitting, and a Hausman-type estimator-disagreement diagnostic, then quantify the detection-damage gap: in the partial-mediation design, where we locate both edges, a band of violations destroys surrogate-index coverage yet is too small to detect on most datasets. On the 64,000-customer Hillstrom experiment the diagnostic rarely flags a violation that biases the surrogate index, and a Cauchy-kernel hybrid of the two estimators inherits 10.7% relative bias; on the 14-million-user Criteo experiment the violation is detected, and subsampling traces detection turning on with scale as damage persists. PPI++ has limits: at a rare-conversion $n = 30,000$ Criteo subsample its empirical coverage is 85.0%. We recommend prespecifying PPI++ with the exact variance as the primary analysis under random labeling and adequate labeled outcome counts, reading the diagnostic as a warning, not a certificate, and treating the surrogate index as a sensitivity analysis.