面向多样化用户行为的排序策略的自适应双重鲁棒离线策略评估
Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior
浏览论文内容
中文总结 AI 辅助
针对多样化用户行为下排序策略的离线评估挑战,提出自适应双重鲁棒(ADR)方法,结合自适应重要加权与奖励回归,在合成实验中降低了均方误差。
中文摘要 AI 辅助
排序策略的离线策略评估(OPE)颇具挑战性,因为从候选集中选择并排序多个物品时,可能的排序数量会随候选物品数量和排序长度呈组合式增长。因此,逆倾向评分(IPS)的重要权重为评估策略与日志策略下的全排序概率之比,可能存在过大的方差。独立逆倾向评分(IIPS)和奖励交互逆倾向评分(RIPS)通过对用户浏览排序的方式施加固定假设来降低方差,但当这些假设与实际行为不匹配时,可能会引入偏差。自适应逆倾向评分(AIPS)通过自适应地对影响每个位置奖励的动作的重要权重进行边缘化,解决了这一权衡问题。当观测到真实用户行为模型时,它在一类无偏的IPS型估计量中达到最小方差。然而,对于更长的排序,其估计精度可能仍会下降,且AIPS未使用奖励模型进行残差校正。我们提出自适应双重鲁棒(ADR)方法,该方法通过控制变量校正将自适应重要加权与奖励回归相结合。我们证明当观测到真实用户行为模型时,ADR具有无偏性,并确定了其相对于AIPS降低方差的充分条件。在每个条件下进行10000次模拟的合成实验中,在一系列日志数据量和排序长度下,ADR相较于AIPS和传统排序OPE估计量,均降低了均方误差。
英文摘要
Off-policy evaluation (OPE) of ranking policies is challenging be- cause selecting and ordering multiple items from a candidate set makes the number of possible rankings grow combinatorially with the number of candidates and the ranking length. Consequently, Inverse Propensity Scoring (IPS), whose importance weight is the full-ranking probability ratio under the evaluation and logging policies, can have excessive variance. Independent IPS (IIPS) and Reward Interaction IPS (RIPS) reduce variance by imposing fixed assumptions on how users browse rankings, but may introduce bias when those assumptions mismatch actual behavior. Adaptive Inverse Propensity Scoring (AIPS) addresses this trade-off by adap- tively marginalizing importance weights over the actions that affect each position-wise reward. It attains minimum variance within a class of unbiased IPS-based estimators when the true user be- havior model is observed. However, its estimation accuracy may still degrade for longer rankings, and AIPS does not use a reward model for residual correction. We propose Adaptive Doubly Robust (ADR), which combines adaptive importance weighting with re- ward regression through a control-variate correction. We establish its unbiasedness when the true user behavior model is observed and characterize a sufficient condition under which it reduces vari- ance relative to AIPS. Across synthetic experiments with 10,000 simulations per condition, ADR improves mean squared error over AIPS and conventional ranking OPE estimators across a range of logged-data sizes and ranking lengths.
发表机构
- Institute of Science Tokyo(东京科学大学)
机构由 AI 辅助整理,请以论文原文为准。