arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于离策略强化学习的大规模生成式检索长期优化

Session-Level Optimization for Large-Scale Retrieval using REINFORCE with Multi-Step Off-Policy Correction

Artem Matveev, Sergei Makeev, Aleksei Krasilnikov, Vladimir Baikalov, Sergei Liamaev, Kirill Khrylchenko

arXiv 2607.02818首次发表:更新:

AI 中文总结

将推荐视为会话级序贯决策问题,用离策略REINFORCE在预收集数据上训练生成式检索器,提出多步重要性权重近似,训练用户反馈模型用于离线评估,引入测试时缩放程序,实验表明方法能改进离线估计。

AI 中文摘要

生成式检索已成为大规模推荐的流行范式。然而,它通常使用监督下一项预测目标进行训练,无法直接优化长期用户满意度。在这项工作中,我们将推荐制定为会话级序贯决策问题,并引入一种自回归方法,用于在预收集的数据上使用离策略REINFORCE训练生成式检索器。与先前工作中使用的单步离策略校正不同,我们提出了一种由自回归公式实现的重要性权重的多步近似。为了支持离线评估,我们训练了一个用户反馈模型,该模型模拟用户对生成的推荐的响应。这使我们能够将会话级序贯决策的双鲁棒离策略评估应用于推荐,这是一个受到有限关注的设置。我们进一步引入了一种基于反馈模型的测试时缩放程序,该程序模拟未来响应并选择具有最高预测长期回报的推荐。在公共大规模Yambda-5B数据集上的实验表明,我们的强化学习智能体在很大程度上保持检索质量的同时,改进了下一项和下一个正预测基线的累积会话奖励的离线估计。此外,在不更新策略的情况下,分配更多推理时间计算来模拟未来响应可以改进基于模型的长期回报估计。

英文摘要

Two-tower models are a widely used paradigm for large-scale retrieval in recommendation. However, they are typically trained with myopic supervised objectives, such as next-item prediction, that do not directly optimize long-term user satisfaction. In this work, we formulate recommendation as a session-level sequential decision-making problem and train a two-tower retriever autoregressively with off-policy REINFORCE on pre-collected data. Unlike the one-step off-policy correction used in prior work, we propose a multi-step approximation of importance weights enabled by the autoregressive formulation. To support offline evaluation, we train a user feedback model that simulates user responses to generated recommendations. This lets us adapt doubly robust off-policy evaluation for sequential decision-making to recommendation, a setting that has received limited attention. We further introduce a feedback-model-based test-time scaling procedure that simulates future responses and selects the recommendation with the highest predicted long-term return. Experiments on the public large-scale Yambda-5B dataset show that our RL agent achieves higher off-policy estimates of cumulative session reward than next-item and next-positive prediction baselines, while remaining competitive on conventional retrieval metrics. Moreover, allocating more inference-time compute to simulating future responses yields higher model-based long-term returns without updating the policy.

CommentsAccepted at the 5th Workshop on End-End Customer Journey Optimization at KDD 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑