arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

方差最优的带联合效应建模的离线策略评估

Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling

Nicolò Felicioni, Michael Benigni, Maurizio Ferrari Dacrema, Paolo Cremonesi

arXiv 2610.08677首次发表:更新:

发表机构

Spotify; Politecnico di Milano(Spotify; 米兰理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出VOCEM估计器,通过插值OffCEM与DR并最小化方差,在23个条件下均优于两基线,实现更稳定的离线策略评估。

AI 中文摘要

离线策略评估(OPE)在上下文赌博机策略中,当动作级别的重要性加权导致过大的方差时变得具有挑战性。双重稳健(DR)估计在共同支持条件下保持无偏性,但保留了这些高方差的动作级别权重。先前的一种估计器——带联合效应模型的离线策略评估(OffCEM),用更稳定的聚类级别权重替换了这些权重,但代价是依赖奖励模型的局部正确性。在本文中,我们表明,在DR和OffCEM所需的假设下,存在一个无偏的估计器族,它在OffCEM和DR之间进行插值。基于这一结果,我们提出了方差最优-CEM(VOCEM)估计器,它选择插值系数以最小化方差。我们以闭式形式推导出总体最优系数,并表明所得估计器的方差不大于任一端点(OffCEM或DR)。在受控合成设置和两个大型动作基准上的实验表明,VOCEM在所有23个评估条件下均优于两个端点,展现出更高的稳定性和经验鲁棒性。

英文摘要

Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance. Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights. A prior estimator, Off-policy evaluation with Conjunct Effect Model (OffCEM), replaces them with more stable cluster-level weights, at the cost of relying on local correctness of the reward model. In this paper, we show that, under the assumptions required by DR and OffCEM, there exists an unbiased family of estimators that interpolates between OffCEM and DR. Building on this result, we propose the Variance Optimal-CEM (VOCEM) estimator, which selects the interpolation coefficient to minimize variance. We derive the population-optimal coefficient in closed form and show that the resulting estimator has variance no larger than either endpoint, OffCEM or DR. Experiments in controlled synthetic settings and on two large-action benchmarks show that VOCEM improves upon both endpoints in all 23 evaluated conditions, exhibiting greater stability and empirical robustness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑