arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36740cs.LG

通过 Top-$k$ 策略分解进行排序策略的高效离线学习

Efficient Offline Learning of Ranking Policies via Top-$k$ Policy Decomposition

发表机构东京科学大学 · 庆应义塾大学 · 康奈尔大学
另 3 家 · 查看机构详情
  • Institute of Science Tokyo(东京科学大学)
  • Keio University(庆应义塾大学)
  • Cornell University(康奈尔大学)
  • Yale University(耶鲁大学)
  • LY Corporation
  • Hanjuku-kaso, Co., Ltd.(Hanjuku-kaso株式会社)

机构由 AI 辅助整理,请以论文原文为准。

Ren Kishimoto, Koichi Tanaka, Haruka Kiyohara, Yusuke Narita, Yasuo Yamamoto, Nobuyuki Shimizu, Yuta Saito

首次发表
浏览论文内容

中文总结 AI 辅助

针对排序策略离线学习中的高方差和偏差问题,提出R-POD方法,将策略分解为top-k选择与底部排序两阶段,结合策略梯度与回归方法,实现高效且无偏的学习。

中文摘要 AI 辅助

许多推荐系统(例如电子商务和新闻平台)旨在为用户提供他们可能与之交互的排名。排序策略的离线策略学习(OPL)使我们能够仅使用历史日志数据来学习新的排序策略。然而,排序设置使 OPL 极具挑战性,因为其动作空间由独特物品的排列组成,规模极其庞大。现有方法主要使用基于策略或基于回归的方法。基于策略的方法通常使用重要性加权策略梯度,可能因动作空间大而遭受高方差问题。另一方面,基于回归的方法使用传统机器学习方法估计期望奖励,避免了方差问题,但可能遭受严重偏差。为了规避现有方法的这些问题,我们提出了一种新的排序 OPL 方法,名为通过 Top-$k$ 策略分解的排序策略优化(R-POD),该方法有效结合了基于策略和基于回归的方法。具体来说,R-POD 将排序策略分解为第一阶段策略(用于选择 top-$k$ 动作)和第二阶段策略(在给定 top-$k$ 动作的情况下选择底部动作)。它使用新的策略梯度估计器学习第一阶段策略,并通过基于回归的方法学习第二阶段策略。该方法可以大幅减少方差,因为它仅对 top-$k$ 动作应用重要性加权。我们还证明,在条件成对正确性条件下,我们针对第一阶段策略的策略梯度估计器是无偏的,该条件仅要求共享相同 top-$k$ 动作的排序对的期望奖励差异可以被正确估计。

英文摘要

Many recommender systems such as for e-commerce and news platforms aim to provide users with rankings they are likely to interact with. Off-Policy Learning (OPL) of ranking policies enables us to learn new ranking policies using only historical logged data. However, ranking settings make OPL remarkably challenging because their action spaces consist of permutations of unique items, being extremely large. Existing methods primarily use either policy- or regression-based approaches. The policy-based approach, which typically uses importance-weighted policy gradients, can suffer from high variance due to large action spaces. The regression-based approach, on the other hand, estimates the expected reward using conventional machine learning methods, avoiding variance issues but potentially suffering from severe bias. To circumvent these issues of existing methods, we propose a new OPL method for ranking, named Ranking Policy Optimization via Top-$k$ Policy Decomposition (R-POD), which combines the policy- and regression-based approaches in an effective fashion. Specifically, R-POD decomposes a ranking policy into a first-stage policy for selecting top-$k$ actions and a second-stage policy for choosing the bottom actions given the top-$k$ actions. It learns the first-stage policy using a new policy gradient estimator and the second-stage policy via the regression-based approach. This method can substantially reduce variance, since it applies importance weighting only to the top-$k$ actions. We also demonstrate that our policy-gradient estimator for the first-stage policy is unbiased under a conditional pairwise correctness condition, which only requires that the expected reward differences of pairs of rankings sharing the same top-$k$ actions can be estimated correctly.

补充信息

↑