基于检验-相关性分解的排名相关检验下推荐排序离线策略评估
Recommendation Ranking Off-Policy Evaluation under Ranking-Dependent Examination via Examination-Relevance Decomposition
浏览论文内容
中文总结 AI 辅助
针对推荐排序离线策略评估中点击偏差问题,提出基于检验-相关性分解的LE-IIPS与ED-DR估计器,后者在检验概率正确或排名无关检验下无偏,大样本时MSE更低。
中文摘要 AI 辅助
离线策略评估利用日志数据估计评估策略的性能,是推荐排序策略的关键。然而,日志点击无法区分未检验项目与已检验但未点击项目,当假设的检验结构不成立时,现有估计器会产生偏差。我们提出两种基于将点击分解为检验与相关性成分的估计器。首先,潜在检验独立逆倾向得分(LE-IIPS)估计器利用策略检验概率比率修正IIPS偏差。其次,检验分解双稳健(ED-DR)估计器将LE-IIPS扩展至双稳健框架。若检验概率正确,无论相关性模型准确性如何,ED-DR均无偏;在排名无关检验下,即使两个模型估计均不准确,ED-DR也无偏。实验表明,在大样本量下,ED-DR相比现有方法实现更低的均方误差(MSE),尤其在检验依赖于排名时。我们同时指出其在样本量较小或级联用户行为条件下的局限性。
英文摘要
Off-policy evaluation, which estimates evaluation policy performance from logged data, is key for recommender ranking policies. However, logged clicks cannot distinguish unexamined items from examined non-clicks, causing bias in existing estimators when the assumed examination structures fail. We propose two estimators based on the decomposition of clicks into examination and relevance. First, the latent-examination independent inverse propensity score (LE-IIPS) estimator corrects the IIPS bias using policy examination probability ratios. Second, the examination-decomposed doubly robust (ED-DR) estimator extends LE-IIPS to a doubly robust framework. ED-DR is unbiased if the examination probabilities are correct regardless of relevance accuracy, or under ranking-independent examination, even if both model estimates are inaccurate. Experiments show that ED-DR achieves a lower MSE than existing methods with large sample sizes, especially when the examination depends on ranking. We also highlight its limitations under small samples or cascade user behavior conditions.
发表机构
- Waseda University(早稻田大学)
机构由 AI 辅助整理,请以论文原文为准。