AI 中文总结
本文梳理A/B测试的三级贝叶斯推断配置,指出贝叶斯因子停止法近最优且可控假阳性率,经验贝叶斯先验可实现错误发现率控制但易受异质因素影响,需根据实验风险选择方法。
AI 中文摘要
A/B测试的贝叶斯推断是一类具有根本不同统计特性的先验和停止规则配置,但常被当作单一方法讨论,目前缺乏系统概述。本文将常见配置组织为三级层级:1)无错误控制的后验一致性;2)通过贝叶斯因子(Bayes factor)停止法在连续监测下限定假阳性率;3)通过经验贝叶斯(empirical Bayes)实现错误发现率控制与校准收缩。许多商业平台默认采用最低层级配置。研究表明,贝叶斯因子停止法对A/B测试文献中多数成本函数而言接近最优,且该规则同时可控制假阳性率,因此决策论与频率论公式的选择主要取决于参数化方式。此外,经验贝叶斯先验是实现第三层级的唯一途径,但获胜者选择的语料库、合并程序及异质指标均会破坏校准,无论语料库规模大小。针对组序贯与始终有效的频率论基线的模拟显示:平坦先验后验停止法会完全复制 naive 窥视(peeking)行为,校准良好的经验贝叶斯先验可实现最低估计误差,仅当推送零效应变体几乎无成本时,期望损失停止法可最小化遗憾。错误率、估计精度与遗憾均为不同风险,合适的方法取决于实验项目需控制的风险,而非反之。
英文摘要
Bayesian inference for A/B testing is a family of prior and stopping-rule configurations with fundamentally different statistical properties, but it is often discussed as a single method, and no systematic overview exists. This paper organizes common configurations into a three-tier hierarchy: 1) posterior coherence with no error control, 2) false positive rates bounded under continuous monitoring via Bayes factor stopping, and 3) false discovery rate control and calibrated shrinkage via empirical Bayes. Many commercial platforms operate at the lowest tier by default. We show that Bayes factor stopping is near-optimal for a broad class of cost functions, including most proposed in the A/B testing literature; because the same rule also controls the false positive rate, the choice between a decision-theoretic and a frequentist formulation is largely one of parameterization. Furthermore, the empirical Bayes prior is the only path to the third tier, but winner-selected corpora, pooled programs, and heterogeneous metrics can each prevent calibration regardless of corpus size. Simulations against group-sequential and always-valid frequentist baselines show that flat-prior posterior stopping exactly reproduces naive peeking, that a well-calibrated empirical Bayes prior achieves the lowest estimation error, and that expected-loss stopping minimizes regret only when shipping a null-effect variant is nearly free. Error rates, estimation accuracy, and regret are all different risks, and the appropriate method follows from the risks an experimentation program needs to control, not the other way around.