arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RecoReward:用于推荐的推荐器引导多模态描述生成

RecoReward: Recommender-Guided Multimodal Description Generation for Recommendation

Guohong Mu, Yueyang Liu, Jiangxia Cao, Changxin Lao, Zijie Zhuang, Yuhui Zhang, Jiaqi Feng, Ruochen Yang, Shuang Yang, Zhaojie Liu, Qibin Hou

arXiv 2607.25901首次发表:更新:

AI 中文总结

研究针对多模态推荐中传统方法不足,提出RecoReward,训练时用行为衍生奖励,保留仅内容推理。在直播推荐中利用用户历史等估计亲和力,通过推荐器亲和力分数提供反馈,实验表明该方法能提升MLLM性能,利于下游推荐且保留内容服务。

AI 中文摘要

多模态大语言模型(MLLMs)可将多模态商品内容转换为结构化描述用作推荐的语义特征。传统仅内容生成无法利用下游用户信号确定应强调的语义。近期用户条件方法通过用户历史或档案纳入这些信号,但推理时需用户信息且生成依赖用户。本文介绍RecoReward,其在训练期间使用行为衍生奖励并保留仅内容推理。在直播推荐中,将历史参与用户作为未来目标用户的代理,用观察到的非目标用户估计用户间广泛共享的亲和力。推荐器亲和力分数(RAS)对比这些信号为强化学习提供用户选择性反馈,使学习到的策略无需用户输入就能生成单一共享描述。离线基准测试中,RecoReward - 9B在七个召回指标上优于其Qwen3.5 - 9B基线及所有其他评估模型。在线A/B测试也显示性能提升。这些结果表明RecoReward训练MLLM生成有益于下游推荐的商品特征,同时保留仅内容服务。

英文摘要

Multimodal large language models (MLLMs) can convert multimodal item content into structured descriptions used as semantic features for recommendation. Conventional content-only generation, however, cannot use downstream user signals to determine which semantics should be emphasized. Recent user-conditioned methods incorporate these signals through user histories or profiles, but they require user information at inference and make generation user-dependent. In this paper, we introduce RecoReward, which instead uses behavior-derived rewards during training and preserves content-only inference. To instantiate this idea in live-stream recommendation, we treat historically engaged users as a proxy for future target users and use observational non-target users to estimate affinity shared broadly across users. The Recommender Affinity Score (RAS) contrasts these signals to provide user-selective feedback for reinforcement learning, allowing the learned policy to generate a single shared description without user inputs. In our offline benchmark, RecoReward-9B outperforms its Qwen3.5-9B baseline and all other evaluated models across seven recall metrics. Online A/B testing also shows performance gains. These results show that RecoReward trains the MLLM to produce item features that benefit downstream recommendation while retaining content-only serving.

Comments16 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑