arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30273cs.LG

离线策略评估作为设计自适应实验的决策支持工具

Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments

João Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出利用离线策略评估与暖启动模拟相结合的方法,基于历史A/B测试数据评估自适应策略的部署价值,并验证其在异质性存在时优于固定分配。研究通过合成实验和开放基准验证了方法的有效性,为自适应实验的决策提供了实用工具。

中文摘要 AI 辅助

我们研究了如何利用固定随机实验(A/B测试)的历史数据来指导基于上下文赌博机的自适应实验的部署。给定在静态分配下收集的数据,我们的目标是评估哪些自适应策略(如果有的话)会优于原始设计,以及在什么条件下会优于原始设计。为此,我们将离线策略评估(OPE)与受控的暖启动模拟相结合。从表现出异质性处理效应的已记录A/B测试数据中,我们估计干扰分量,并使用双重稳健估计器对一组预先指定的自适应和非自适应策略进行排序。当真实值可用时,我们随后在模拟器中部署相同的离线训练策略,该模拟器重用精确的数据生成奖励概率,提供了一个安全的、以真实值为锚定的环境,用于研究在暖启动下从离线到在线的过渡。使用具有已知异质性结构和预言策略的合成随机对照试验,我们的结果表明,当存在有意义的异质性时,自适应、上下文感知的策略会优于固定分配,而在没有异质性时则几乎没有益处。我们在标准开放基准(Hillstrom、Criteo Uplift和LaLonde)上加强了我们的发现,并通过策略价值和遗憾视角进行重新解读。总体而言,我们的结果提供了一种实用的方法论,用于决定何时值得部署自适应实验,以及如何利用现有的A/B测试数据在竞争的自适应策略中进行选择。

英文摘要

We investigate how historical data from fixed randomized experiments (A/B tests) can be used to inform the deployment of adaptive experiments based on contextual bandits. Given data collected under a static allocation, our goal is to assess which adaptive policies, if any, would have outperformed the original design and under what conditions. To this end, we combine off-policy evaluation (OPE) with a controlled warm-start simulation. From logged A/B test data exhibiting heterogeneous treatment effects, we estimate nuisance components and use doubly robust estimators to rank a portfolio of pre-specified adaptive and non-adaptive policies. When ground truth is available, we then deploy the same offline-trained policies in a simulator that reuses the exact data-generating reward probabilities, providing a safe, ground-truth-anchored environment to study the offline-to-online transition under warm starting. Using synthetic randomized controlled trials with known heterogeneity structures and an oracle policy, our results indicate that adaptive, context-aware policies improve upon fixed allocations when meaningful heterogeneity is present, while providing little benefit in its absence. We reinforce our findings on standard open benchmarks (Hillstrom, Criteo Uplift, and LaLonde), reinterpreted through a policy-value and regret perspective. Overall, our results provide a practical methodology for deciding when adaptive experimentation is worth deploying and how to select among competing adaptive policies using existing A/B test data.

补充信息

↑