arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10167stat.MEcs.LGmath.STstat.TH

弱信号下的策略学习

Policy Learning with Weak Signals

Benedikt Koch, Winston Chou, Aurélien Bibaut, Nathan Kallus

首次发表
浏览论文内容

中文总结 AI 辅助

针对数字实验中弱信号、高维协变量和海量数据的挑战,证明一般情形下最优策略不可学习,但在处理效应平滑时提出基于线性平滑器的极小极大自适应策略,并在Netflix真实实验中验证其优于非个性化策略。

中文摘要 AI 辅助

数字实验中的策略学习面临三大挑战:弱的信噪比、丰富的协变量空间以及海量的数据规模。我们通过将来自日益精细的协变量分区的处理效应估计建模为具有有界信噪比的高斯观测,来形式化这一情形。我们证明,在一般情况下,最优处理策略在此设定下是不可学习的。即使学习最优策略值也面临不切实际的缓慢速率。然而,当处理效应平滑变化时,我们基于线性平滑器推导出极小极大自适应策略,这些策略实现了趋于零的福利遗憾。我们通过将我们的框架应用于Netflix的大规模真实世界实验来展示其实用价值,表明即使在具有挑战性的实证环境中,个性化线性平滑策略也能优于非个性化策略。

英文摘要

Policy learning in digital experimentation faces three challenges: weak signal-to-noise ratios, rich covariate spaces, and massive data volumes. We formalize this regime by modeling treatment-effect estimates from increasingly fine covariate partitions as Gaussian observations with bounded signal-to-noise ratios. We establish that, in general, the optimal treatment policy is not learnable in this setting. Even learning the optimal policy value suffers from impractically slow rates. However, when treatment effects vary smoothly, we derive minimax-adaptive policies based on linear smoothers that achieve vanishing welfare regret. We demonstrate the practical value of our framework by applying it to large-scale real-world experiments at Netflix, showing that personalized linear-smoothing policies can dominate unpersonalized policies even in this challenging empirical setting.

发表机构

  • Harvard University(哈佛大学)
  • Netflix(奈飞)

机构由 AI 辅助整理,请以论文原文为准。

↑