arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2507.21274cs.LG

面向多样化与新颖性推荐的大语言模型增强强化学习

Large Language Model-Enhanced Reinforcement Learning for Diverse and Novel Recommendations

  • Carnegie Mellon University(卡内基梅隆大学)
  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

Jiin Woo, Alireza Bagheri Garakani, Tianchen Zhou, Zhishen Huang, Yan Gao

更新

AI总结:

针对推荐系统多样性不足、强化学习随机探索与用户兴趣不匹配的问题,提出LAAC方法,以LLM为参考策略推荐新颖物品,结合双层优化训练轻量策略,通过正则化缓解高估问题,在多指标上优于基线且无需昂贵微调。

AI中文摘要:

在推荐系统中,多样性与新颖性对捕捉多元用户偏好、鼓励探索至关重要,但多数系统更侧重点击相关性。尽管已有研究探索用强化学习(RL)提升多样性,但这类方法常依赖随机探索,可能与用户兴趣不匹配。我们提出LAAC(LLM-guided Adversarial Actor Critic,大语言模型引导的对抗式演员-评论家),这是一种新方法,以大语言模型(LLMs)作为参考策略推荐新颖物品,同时训练一个轻量策略,利用系统特定数据优化这些推荐。该方法将训练构建为演员网络与评论家网络之间的双层优化,使评论机能选择性地支持有潜力的新颖动作,演员则能在LLM推荐的基础上进一步优化策略。为缓解对不可靠LLM推荐的高估问题,我们应用正则化,将未探索物品的评论家值锚定在估计充分的数据集动作附近。在真实世界数据集上的实验表明,LAAC在多样性、新颖性和准确性上均优于现有基线,且在不平衡数据上保持鲁棒性,无需昂贵的微调即可有效整合LLM知识。

英文摘要:

In recommendation systems, diversity and novelty are essential for capturing varied user preferences and encouraging exploration, yet many systems prioritize click relevance. While reinforcement learning (RL) has been explored to improve diversity, it often depends on random exploration that may not align with user interests. We propose LAAC (LLM-guided Adversarial Actor Critic), a novel method that leverages large language models (LLMs) as reference policies to suggest novel items, while training a lightweight policy to refine these suggestions using system-specific data. The method formulates training as a bilevel optimization between actor and critic networks, enabling the critic to selectively favor promising novel actions and the actor to improve its policy beyond LLM recommendations. To mitigate overestimation of unreliable LLM suggestions, we apply regularization that anchors critic values for unexplored items close to well-estimated dataset actions. Experiments on real-world datasets show that LAAC outperforms existing baselines in diversity, novelty, and accuracy, while remaining robust on imbalanced data, effectively integrating LLM knowledge without expensive fine-tuning.

↑