arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过离线评估方法的A/B测试实现更准确的算法比较

A More Accurate Algorithm Comparison through A/B Testing using Offline Evaluation Methods

Koki Konishi, Masataka Ushiku, Yuta Saito

arXiv 2607.01958首次发表:更新:

发表机构

Hakuhodo DY Holdings Inc.; Cornell University(博报堂DY控股公司; 康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

揭示A/B测试可能比离线评估产生更高算法选择错误率的反直觉现象,提出通过引入假设中间算法并逐步估计性能差异来诱导正相关,从而减少关键选择错误。

AI 中文摘要

A/B测试是在线服务中选择更优算法的黄金标准。由于A/B测试的实验成本高且存在降低用户体验和收入的风险,离线评估作为更安全的替代方案受到关注,但公认离线评估的估计精度显著较低。因此,最终选择决策通常通过A/B测试进行。与传统观点相反,我们揭示了一个反直觉现象:A/B测试可能比离线评估产生更高的算法选择错误率。这是因为A/B测试中使用的样本均值估计器不会诱导正相关,而正相关对于减少关键选择错误(即低估真正更优算法和高估真正更劣算法)至关重要。相比之下,离线评估方法在估计和比较多个算法的性能时,依赖共享的离线数据,无意中产生了这种有益的相关性。基于这一见解,我们提出了一种在A/B测试中通过有意诱导正相关来改进算法选择的估计器。关键思想是引入一个假设的中间算法,并在每一步使用共享数据逐步估计算法A、M和B之间的性能差异。这种方法使得每一步都能应用离线评估技术,从而诱导正相关并减少关键选择错误。此外,我们推导了关于结果方差的最优中间算法,并通过偏差-方差分析分析了其相对于现有方法的优势。在真实数据上的实验表明,我们的估计器在使用仅一半A/B测试数据的情况下,达到了与现有方法相同的选择错误率。

英文摘要

A/B testing is the gold standard for selecting the better algorithm in online services. While offline evaluation has attracted attention as a safer alternative due to the high experimental costs and the potential risk of degrading user experience and revenue in A/B testing, it is widely recognized that the estimation accuracy of offline evaluation is substantially lower. As a result, final selection decisions are typically made through A/B testing. Contrary to this conventional view, we reveal a counterintuitive phenomenon in which A/B testing can produce a higher algorithm selection error rate than offline evaluation. This occurs because the sample mean estimator used in A/B testing does not induce positive correlation, which is crucial for reducing critical selection errors, namely underestimating the truly superior algorithm and overestimating the truly inferior one. In contrast, offline evaluation methods unintentionally generate this beneficial correlation by relying on shared offline data when estimating and comparing the performance of multiple algorithms. Building on this insight, we propose an estimator that intentionally induces positive correlation to improve algorithm selection in A/B testing. The key idea is to introduce a hypothetical middle algorithm and to estimate the performance difference between algorithms A, M, and B in a stepwise manner using shared data at each step. This approach enables the application of offline evaluation techniques in each step, thereby inducing positive correlation and reducing critical selection errors. Furthermore, we derive the optimal middle algorithm regarding the resulting variance and analyze its advantages over existing methods through bias-variance analysis. Experiments on real-world data demonstrate that our estimator achieves the same selection error rate as existing approaches while using only one half of the A/B testing data.

Comments13 pages, 10 figures. Accepted at KDD 2026; this version extends the camera-ready with additional experiments in Appendix B

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑