arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当离线评估产生误导时:延迟反馈情境多臂老虎机中奖励与策略选择的诊断协议

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan

arXiv 2608.11560首次发表:更新:

发表机构

Thumbtack, Inc.(Thumbtack公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对延迟反馈情境多臂老虎机的奖励与策略选择,提出有序诊断协议,验证其有效性并揭示奖励排序易误导、个性化溢价易高估的结论。

AI 中文摘要

利用情境多臂老虎机(CMAB)定制营销信息可带来实际商业价值,但最终关键目标——下游转化——需数周后才能观测到,此时已无法用于在线学习。因此团队需基于快速代理奖励训练老虎机,同时还需判断情境多臂老虎机是否值得因其复杂性而优于发送单一最佳信息。采用常规离线检查方法——批量离策略估计、边际臂区分测试、置信区间——在延迟反馈场景下可能产生系统性误导。本文提出一种有序诊断协议,在信任任何报告的提升值前,从两个维度筛选奖励与策略候选:对齐性(优化奖励是否能推动核心目标)和可学习性(老虎机能否识别奖励最优策略)。我们在已知真实值的场景中验证了该协议——包括公开离策略评估基准和可控合成生成器,并在已部署的大型市场推送系统中进行了演示(该系统含5个臂和1个拆分,证据具有方向性而非统计效力)。得出两个核心结论:(N1)单一离线数值可能错误排序奖励:更密集的奖励信号能为老虎机提供更多学习内容,因此在静态估计中看似持平的奖励,在线学习时会产生差异;(N2)若无法提前确定单一最佳信息,基于用户的策略部分只是避免押注错误选项——这看似个性化实则是鲁棒性,因此“个性化溢价”易被高估。本文的贡献是方法论层面而非算法层面:有序诊断协议、其揭示的两个核心结论,以及将其应用于延迟反馈情境多臂老虎机的端到端经验。

英文摘要

Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message. Settling both decisions with the usual offline checks - a batch off-policy estimate, a marginal arm-discrimination test, a confidence interval - can mislead systematically under delayed feedback. We give an ordered diagnostic protocol that screens a reward-and-policy candidate on two axes, alignment (does optimizing the reward move the north-star?) and learnability (can the bandit identify the reward-optimal policy?), before trusting any reported lift. We validate it where the truth is known - a public off-policy-evaluation benchmark and a controllable synthetic generator - and illustrate it on a deployed large-marketplace push system (where, with five arms and one split, the evidence is directional rather than powered). Two lessons recur. (N1) A single offline number can mis-rank rewards: a denser reward signal gives the bandit more to learn from, so rewards that look tied in a static estimate pull apart once learning happens online. (N2) If you cannot tell in advance which single message is best, a per-user policy partly just avoids betting on the wrong one - that looks like personalization but is really robustness, so a "personalization premium" is easily overstated. Our contribution is methodological rather than algorithmic: the ordered protocol, the two lessons it surfaces, and the end-to-end experience of applying it to a delayed-feedback CMAB.

CommentsAccepted at the 5th Workshop on End-to-End Customer Journey Optimization (KDD 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑