稳健的A/B决策
Robust A/B Decisions
浏览论文内容
中文总结 AI 辅助
针对A/B测试中传统显著性检验的缺陷,提出基于模糊厌恶的稳健决策规则,利用Donsker-Varadhan表示实现简单封闭形式,并在552个广告实验中显著降低遗憾。
中文摘要 AI 辅助
A/B测试是公司决策中的标准做法。在标准流程中,实验数据通过应用均值差异(提升度)的t检验转换为部署决策,如果提升度为正值且具有统计显著性,则部署处理方案。这种常见的工作流程回答了一个错误的问题。我们认为,公司需要的是针对未来部署环境中经济收益的决策规则,而不是对实验样本中相等性的检验。我们开发了一个模糊厌恶决策框架,在该框架中,每个臂通过其在接近实验结果分布的分布上的模糊惩罚值进行评估。得益于Donsker-Varadhan表示,所得规则具有简单的封闭形式,并且仅需要标准A/B测试的结果数据加上一个控制实验信任度的可解释参数。因此,我们的规则在实施难度上不高于t检验。均值-方差近似展示了该规则如何惩罚变异性,而与效用最大化的联系则表明它是一个确定性等价。我们能够利用来自美国匿名在线平台的552个广告实验档案,在数字营销背景下对我们提出的规则进行真实世界评估。与传统的假设检验相比,所提出的规则大幅减少了遗憾。结果表明,经济上保守、分布感知的部署规则在数字实验中可以优于统计显著性规则。
英文摘要
A/B tests are standard in firm decision making. In the standard pipeline, experimental data is converted to a deployment decision by applying a t-test of the difference in means (the lift) and deploying the treatment if lift is positive and statistically significant. This common workflow answers the wrong question. We argue that firms need a decision rule for economic payoffs in the future deployment environment, not a test of equality in the experimental sample. We develop an ambiguity-averse decision framework in which each arm is evaluated by its ambiguity-penalized value over distributions close to the experimental outcome distribution. The resulting rule has a simple closed form thanks to the Donsker-Varadhan representation and it requires only the outcome data from a standard A/B test plus one interpretable parameter governing trust in the experiment. Our rule is thus no more difficult to implement than a t-test. A mean-variance approximation shows how the rule penalizes variability, while a connection to utility maximization shows it to be a certainty equivalent. We are able to perform a real-world evaluation of our proposed rule in the context of digital marketing using an archive of 552 advertising experiments from an anonymous US-based online platform. The proposed rule substantially reduces regret relative to conventional hypothesis testing. The results show that economically conservative, distribution-aware deployment rules can outperform statistical-significance rules in digital experimentation.
发表机构
- UC Santa Barbara(加州大学圣塔芭芭拉分校)
- Graduate School of Business, Stanford University(斯坦福大学商学院)
- Booth School of Business, University of Chicago(芝加哥大学布斯商学院)
机构由 AI 辅助整理,请以论文原文为准。