arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为基于MLLM的推荐系统扩展可表述理由

Scaling Articulated Rationales for MLLM-based Recommendation

Haoke Xiao, Yueyang Liu, Yuhui Zhang, Xiang Chen, Yufei Liu, Jia Xu, Yalong Guan, Xiaolan Zhu, Xiaoyu Zhang, Shijun Wang, Shuang Yang, Zijie Meng, Zejian Zhang, Ruochen Yang, Xiangyu Wu, Tingting Gao, Han Li, Lantao Hu, Cheng Luo, Kun Gai

arXiv 2609.17639首次发表:更新:

发表机构

Kuaishou Technology(快手科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出SARA框架,将稀疏的用户可表述理由扩展为工业推荐系统的可扩展信号,通过数据引擎、MLLM对齐和排序集成,提升参与度并减少负面反馈。

AI 中文摘要

现代推荐系统主要从点击、观看时长和负面反馈等隐式行为中推断用户偏好,但这些信号揭示了用户的行为,而非他们喜欢或不喜欢内容的原因。本研究将用户可表述理由(AURs),即用户对其偏好的自然语言解释,作为一类新的极性感知和理由级文本信号用于推荐。尽管AURs具有潜在价值,但由于其天然稀疏、质量往往较低且仅覆盖一小部分物品,难以在工业系统中使用。我们提出了SARA(扩展可表述理由),一个将稀疏AURs转化为可扩展推荐信号的工业框架。SARA首先构建了一个数据引擎,从2.4亿快手直播用户中获取并策展AURs,生成SARA-HQ,一个质量受控且以作者为中心的理由数据集。然后,通过大规模SFT和质量精炼DPO,将通用MLLM对齐为SARA-7B,将理由生成从86,564名覆盖AUR的作者扩展到全1000万作者空间。最后,SARA-Ranker通过理由感知交互建模和拒绝记忆建模,将生成的正向和负向理由集成到生产排序中。广泛的离线评估、人工校准和在线A/B测试表明,SARA-7B生成的比强MLLM基线更具体、极性一致且有依据的理由,而SARA-Ranker在生产中提升了参与度并减少了负面反馈。SARA已部署并每日刷新超过30天,确立了可表述理由作为工业推荐系统实用的一等文本信号的地位。

英文摘要

We presented SARA, an industrial framework that transforms sparse articulated user rationales into scalable recommendation signals. Its data engine curates questionnaire responses into SARA-HQ, providing explicit preference supervision for aligning SARA-7B through SFT and Quality-Refining DPO. This alignment extends rationale generation from $86{,}564$ questionnaire-covered authors to the full $10$M-author space. SARA-Ranker translates the generated positive and negative rationales into features for user--author interaction modeling and negative-feedback history modeling, connecting articulated reasons to production ranking. Evaluation on unseen authors demonstrates that SARA-7B generates more specific, relevant, and grounded rationales than the evaluated general-purpose MLLMs. On top of a strong industrial ranking baseline with multimodal features, separate online A/B tests show that positive-rationale integration increases watch time by $0.99\%$, while negative-rationale integration reduces Hate feedback by $8.16\%$. Daily refresh and more than $30$ days of production deployment further demonstrate the operational feasibility of the approach. These findings establish articulated rationales as a useful complement to behavioral and content signals, and demonstrate a practical role for MLLMs in scaling sparse human explanations into preference information that improves industrial recommendation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑