arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过反事实估计加速A/B测试:通过策略重叠降低方差

Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap

Olivier Jeunen

arXiv 2607.14604首次发表:更新:

AI 中文总结

研究在线控制实验中A/B测试方差大成本高问题,提出利用策略重叠加速实验的新协议,将随机处理分配机制视为元策略,用Δ-离策略估计方法获无偏估计,理论和实证证明该方法可提升效率,有望用于多种系统评估。

AI 中文摘要

在线控制实验是在线平台假设检验的黄金标准。尽管其无处不在,但运行成本高昂,方差问题影响治疗效果评估的统计效力。标准方差减少技术利用基于模型的控制变量减少结果噪声,但对竞争策略间潜在结构关系不了解。本文发现标准A/B测试协议的关键低效率:处理与控制策略在行动上一致时,结果产生噪声而非治疗效果信号,不必要地扩大置信区间。提出利用策略重叠加速实验的新协议,将随机处理分配机制视为元策略,利用Δ-离策略估计方法获得平均治疗效果的无偏估计。理论证明该方法在一般情况下恢复标准A/B测试实践,方差随策略间差异而非原始结果方差缩放。实证结果证实理论见解,有望对推荐系统、信息检索管道和大语言模型接口的实际评估产生重大影响。

英文摘要

Online controlled experiments are the gold standard for hypothesis testing in online platforms. Notwithstanding their ubiquity, they are notoriously expensive to run, and issues of variance hamper statistical power in assessing treatment effects. While standard variance reduction techniques leverage model-based control variates to reduce outcome noise, they remain agnostic to potential structural relationships between competing policies. In this work, we identify a critical inefficiency in the standard A/B-testing protocol: when a treatment and control policy agree on an action, the resulting outcome contributes noise but no signal regarding the treatment effect -- unnecessarily inflating confidence intervals. We propose a novel experimental protocol that exploits this policy overlap to accelerate experimentation. The key insight is to frame the randomised treatment assignment mechanism as a meta-policy, and leverage $Δ$-Off-Policy Estimation methods to obtain unbiased estimates for average treatment effects. We prove analytically that our approach recovers standard A/B-testing practices in the general case, but that its variance scales with the divergence between policies rather than raw outcome variance. Hence, we dominate the standard Difference-in-Means estimator whenever policies have common support, and the improvement is strict whenever the overlap region contributes non-zero residual variance. Empirical results corroborate these theoretical insights -- holding promise for significant impact on the real-world evaluation of recommender systems, information retrieval pipelines, and large language model interfaces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑