arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过模型无关的下游奖励学习实现长期用户参与度优化

Long-term User Engagement Optimization through Model-agnostic Downstream Rewards Learning

Dingsu Wang, Filip Ryzner, Kelly He, Armando Ordorica, David Woo, Aditya Mantha, Liyao Lu, Usha Amrutha Nookala, Haoran Guo, Jiacong He, Olafur Gudmundsson, Matt Chun, Krystal Benitez, Haibin Xie, Alekhya Pyla, Sameer Jain, Zhongjian Jiang, Shruthi Hariharan, Dhruvil Deven Badani, Yijie Dylan Wang

arXiv 2607.14192首次发表:更新:

发表机构

Pinterest(拼趣(图片分享社交平台))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何在推荐系统中优化长期用户价值,提出统一的模型无关下游奖励框架,先制定问题并开发离线筛选框架识别相关行为,再提出奖励信号,经实验验证可提升参与度和留存率,已在多页面部署。

AI 中文摘要

随着推荐系统在过去几年中逐渐成熟,其优化目标已从主要关注短期行为信号转变为更广泛地强调长期用户参与度和留存率。然而,直接优化留存率很困难,因为回报信号稀疏、延迟,且仅部分归因于早期推荐。先前的工作通过序列建模和强化学习解决了这一挑战,但这些方法通常需要特定任务的奖励工程、大量计算开销以及难以推广的表面特定实现。在本文中,我们提出了一个统一的、与模型无关的下游奖励框架,用于在大规模推荐系统中优化长期用户价值。首先,我们制定了下游奖励学习问题,并开发了一个离线筛选框架,以识别早期可观察且能预测未来留存率的会话级行为。然后,我们提出了几个从多个来源观察到的用户行动模式中得出的与模型无关的下游奖励信号。我们进一步讨论了将所提出的奖励推导投入生产的工程工作以及将它们添加到我们的排名模型时所面临的挑战。在线A/B实验表明,与参与度和留存率相关的指标持续改善,并且该框架已部署在多个Pinterest页面上,包括首页、相关Pin、搜索和通知。

英文摘要

As recommender systems mature in the past few years, their optimization objectives have evolved from a primary focusing on short-term behavioral signals to a broader emphasis on long-term user engagement and retention. However, directly optimizing retention is difficult because return signals are sparse, delayed, and only partially attributable to earlier recommendations. Prior work has addressed this challenge with sequential modeling and reinforcement learning, but these approaches typically require task specific reward engineering, substantial computational overhead, and surface specific implementations that are difficult to generalize. In this paper, we present a unified, model-agnostic downstream reward framework for optimizing long-term user value in large-scale recommendation systems. First, we formulate the downstream reward learning problem and develop an offline screening framework to identify session level behaviors that are both observable early and predictive of future retention. We then propose several model-agnostic downstream rewards signals derived from observed user action patterns across multiple sources. We further discuss the engineering effort to productionize the proposed rewards derivations and challenges we faced when adding them to our ranking models. Online A/B experiments demonstrate consistent improvements in engagement and retention-related metrics, and the framework has been deployed across multiple Pinterest surfaces, including Homefeed, Related Pins, Search, and Notifications.

CommentsRecsys 2026, revised version

DOI:10.1145/3773078.3831912

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑