arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08043stat.AP

长期评估:来自业界的经验教训

Evaluating for the long term: Learnings from industry

Leif Sigerson, Tom Cunningham, Winston Chou, Sana Pandey, Jonathan Stray, Lo-Hua Yuan, Eytan Bakshy, Timothy Chan, Molly Davies, Maria Dimakopoulou, Simon Ejdem… 展开作者

Leif Sigerson, Tom Cunningham, Winston Chou, Sana Pandey, Jonathan Stray, Lo-Hua Yuan, Eytan Bakshy, Timothy Chan, Molly Davies, Maria Dimakopoulou, Simon Ejdemyr, Kenneth Hung, Nathan Kallus, Madhav Kumar, Thu Le, A. Demetri Pananos, Lee Richardson, Brennan Schaffner, Rose Tan, Martin Tingley, Nadia Tomova, Panagiotis Toulis, Wenjing Zheng, Zander Arnao, Dean Eckles

首次发表
浏览论文内容

中文总结 AI 辅助

本文结合26位专家的研讨成果,提出短期实验处理效应符号反转罕见,单变量自动代理难超越,简单可解释代理更优,实验学习代理需长期实验组合,且无法替代优质长期实验。

中文摘要 AI 辅助

在线平台优先考虑长期业务成果,但典型的实验时长过短,无法直接衡量这些成果。本文的目标是收集并分享业界知识,内容涉及如何从短期实验中做出更符合长期成果的决策。基于一场为期一天的研讨会,参会者包括来自15家在线平台和4所大学的26位专家,我们提出了一系列反映当前业界知识的命题。参会者普遍认为,短期到长期处理效应的符号反转情况较为罕见,反转集中在特定场景,例如涉及内容质量信号、过度货币化以及定价的处理。尽管处理效应的幅度会随时间变化,但对应于短期处理对感兴趣的长期指标的效应的“单变量自动代理”通常很难被超越。一个反复出现的主题是,代理的重要性不仅在于(甚至主要不在于)对真实长期结果无偏,而在于能改善决策。因此,参会者普遍认为,简单、可解释的代理通常比复杂但难以解释的代理指标更可取。参会者还一致认为,由于对混淆和可迁移性的担忧,从实验中学习的代理通常比从观察中学习的代理更可取。然而,缺点在于,从实验中学习优质代理通常需要大量、具有代表性的长期实验组合,而很少有平台具备这样的条件。我们的结论是,无论是学习代理还是验证代理,都无法替代运行良好的长期实验,并且我们强调了未解决的挑战,包括不断演变的处理、未完全由短期代理介导的持续处理,以及实验样本与目标总体之间的不匹配。

英文摘要

Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and share industry knowledge on how to make decisions from short-term experiments that are better aligned with long-term outcomes. Based on a daylong workshop with 26 experts from 15 online platforms and 4 universities, we formulate a series of propositions that reflect current industry knowledge. Participants largely agreed that reversals of sign from short-run to long-run treatment effects are rare, with reversals concentrating in specific cases such as treatments involving content quality signals, hyper-monetization, and pricing. Although the magnitude of treatment effects can shift over time, a "univariate autosurrogate", corresponding to the short-run treatment effect on the long-run metric of interest, is often hard to beat. A recurring theme was the importance of surrogates that are not only (or even primarily) unbiased for true long-run outcomes, but that improve decision-making. Thus, participants generally agreed that simple, interpretable surrogates were generally preferable to elaborate but hard-to-explain surrogate indices. Participants also agreed that, due to concerns about confounding and transportability, experimentally-learned surrogates are generally preferable to observationally-learned surrogates. However, the drawback is that learning good surrogates from experiments typically requires a large, representative portfolio of long-run experiments that few platforms possess. We conclude that there is no substitute for a well-run long-term experiment, whether for learning surrogates or validating them, and we highlight open challenges including evolving treatments, persistent treatments not fully mediated by short-term proxies, and mismatch between experimental samples and the target population.

补充信息

↑