arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

安全作为约束:微调LLM推荐系统以自我解释

Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself

Jiashu He, Emma Yanyang Kong, JJ Tan, David Fagnan

arXiv 2609.13657首次发表:更新:

发表机构

University of Pennsylvania; Netflix(宾夕法尼亚大学; 奈飞)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过受约束GRPO微调推荐LLM,在保持推荐性能的同时生成忠实且无害的个性化解释,使三项标准通过率从0.649提升至0.956。

AI 中文摘要

传统推荐系统通常被训练用于预测用户接下来会与哪个项目互动,但不预测原因。然而,为用户可能喜欢预测项目提供个性化证据,是增强服务并提高用户对推荐真正感兴趣可能性的重要方式。这一服务可以通过将前沿模型调用集成到面向会员的流程中来提供,但这会增加额外成本和延迟。在本文中,我们训练了一个推荐LLM,基于大型视频流服务中用户的观看历史,为其推荐生成个性化解释。我们对生成的解释施加了两个要求:它必须忠实于其所关联的节目元素,并且必须对用户严格无害。为此,我们首先训练了两个LLM评判奖励模型,涵盖三个具体标准,并提出了受约束的GRPO以整合这些不同标准。在一个留出的真实世界测试集上,我们微调的模型在自身评判下,三项标准全部通过的比率从0.649提升至0.956,在独立评判下从0.677提升至0.931,而前沿生成器的表现与未调优的推荐基线相似。我们进行了进一步实验,以表明模型的语言和推荐能力保持不变。基于这些结果,我们得出结论:基于LLM的推荐系统可以在不损害其原始推荐性能的情况下,针对其他复杂任务进行微调,从而为由单一模型驱动的进一步智能体用户界面提供见解。

英文摘要

Traditional recommender systems are typically trained to predict what item users will interact with next, but not why. However, offering personalized evidence for why a user might like the predicted item is an important way to enhance the service and to raise the likelihood that the user will be genuinely interested in the recommendation. This service can be delivered by integrating a frontier-model call into the member-facing pipeline, but it will add extra cost and latency. In this paper, we train a recommender LLM to generate personalized explanations for its reccomendation, based on the user's watching history at a large video streaming service. We impose two requirements on the generated explanation: it must be faithful to the elements of the shows it links, and it must be strictly non-harmful to the user. To this end, we first train two LLM-judge reward models covering three specific criteria, and propose constrained GRPO to incorporate these different criteria. On a held-out real-world testing set, our fine-tuned model improves the all-three-criteria PASS rate rises from 0.649 to 0.956 under our own judges and from 0.677 to 0.931 under an independent judge, where as the frontier generator performs similar to the untuned recommender baseline. We conduct further experiments to show that the model's language and recommendation abilities remain unchanged. Based on these results, we conclude that an LLM-based recommender can be fine-tuned on other complex tasks without compromising its original recommendation performance, thus provide insights for further agentic user interface powered by a single model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑