arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RECAP:用于短视频推荐的反馈驱动流语义用户画像

RECAP: Feedback-Driven Streaming Semantic User Profiles for Short-Video Recommendation

Ziyi Zhao, Xiaoyou Zhou, Xiao Lv, Yangyang Li, Chubo He, Zhao Liu, Jiayao Shen, Yuqi Liu, He Li, Chengyi Zhang, Jian Liang, Ming Li, Chongming Gao, Fuli Feng, Ruiming Tang, Han Li

arXiv 2607.15730首次发表:更新:

AI 中文总结

研究短视频推荐中用户画像生成问题,提出RECAP离线闭环框架,结合大语言模型语义更新、生命周期及容量控制维护画像,通过大语言模型判断和双塔评估器构建反馈,实验证明其提升推荐指标,在线测试显示提高用户应用使用时间。

AI 中文摘要

基于语言的用户画像将长期行为历史转化为用于推荐的显式语义表示。然而,大多数画像生成器是开环优化的,在现实世界的短视频推荐中,用户行为如流到达,画像必须在有限容量下增量更新。我们提出RECAP,一个离线闭环框架,通过历史推荐反馈优化流结构化语义画像。它通过结合基于大语言模型的语义更新与确定性生命周期和容量控制来维护画像,通过大语言模型判断过滤标签一致行为对并训练双塔评估器构建画像目标语义反馈。实验表明RECAP比基础生成器提升了uAUC和Recall@2000,在线A/B测试显示平均应用使用时间有显著提升。

英文摘要

Language-based user profiles convert long behavioral histories into explicit semantic representations for recommendation. However, most profile generators are optimized in an open loop: they may summarize past behavior fluently, but are not directly trained to improve future recommendation. We study this problem in real-world short-video recommendation, where user behaviors continuously arrive as streams and profiles must be incrementally updated under limited capacity. This requires maintaining a consistent bounded profile state and constructing profile-targeted semantic feedback from industrial implicit behavior logs. We propose RECAP, an offline closed-loop framework for optimizing streaming structured semantic profiles with historical recommendation feedback. RECAP maintains each profile as a bounded structured memory by combining LLM-based semantic updates with deterministic lifecycle and capacity control. RECAP constructs profile-targeted semantic feedback by filtering label-consistent behavior pairs with an LLM judge and training a dual-tower evaluator whose matching score serves as a GRPO reward. Experiments on Kuaishou short-video data show that RECAP improves uAUC by 0.0084 and Recall@2000 by about 4.9% over the base generator. Further analyses confirm the benefits of feedback construction and policy optimization, and show more grounded refinement and user-level abstraction in profile updates. A seven-day online A/B test further shows a statistically significant 0.139% improvement in average application usage time per user.

Comments9 pages, 5 figures, 2 tables. Accepted at RecSys 2026

DOI:10.1145/3773078.3831930

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑