arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于记忆增强推理的大语言模型智能体群组推荐增强方法

Enhancing Group Recommendation with Memory-Augmented Reasoning in LLM Agent

Qimeng Niu, Bowen Hao, Zixuan Zhang, Shuyu Qu, Hongzhi Yin

arXiv 2608.21939首次发表:更新:

AI 中文总结

本研究提出基于LLM智能体的AGR模型,通过记忆模块与推理模块建模群组偏好动态演变及决策过程,在公开数据集上显著提升了群组推荐的准确率与可解释性。

AI 中文摘要

群组推荐的核心挑战在于建模用户偏好的动态演变以及解释共识形成过程。现有基于大语言模型(LLM)的方法虽提升了可解释性,但将交互历史视为固定文本,忽略群组/用户偏好随时间的自然演变,且缺乏对复杂群组决策过程的显式建模。为解决这些问题,我们提出AGR,一种基于LLM的智能体,由记忆模块与推理模块构成。记忆模块采用基于令牌的哈希表动态管理群组与用户的历史交互,支持插入、更新、检索、遗忘无关记录及汇总演变的群组与用户画像以实现高效追踪等基础操作。推理模块基于检索到的动态画像执行多步推理,包括群组兴趣收集、群组共识优化、多维度评估及可解释推荐生成,从而突破黑箱推理,提供完全可解释的推荐。实践中,我们采用强化微调(RFT)范式:首先用监督微调(SFT)使模型具备调用记忆与推理模块的基础能力,再用群组相对策略优化(GRPO)提升其协调这些模块的自主能力。在LastFM与Douban数据集上的实验表明,AGR在推荐准确率与可解释性上均显著优于现有最优方法。我们的模型已开源至该httpsURL。

英文摘要

The core challenge in group recommendation lies in modeling the dynamic evolution of user preferences and explain?ing the consensus formation process. Existing Large Language Model (LLM)-based methods, despite improved interpretability, treat interaction history as fixed text, ignoring the natural evolution of group/user preferences over time, and lacking explicit modeling of the complex group decision-making process. To address these issues, we propose AGR, a LLM-based agent, which consists of a Memory Module and a Reasoning Module. The Memory Module employs a token-based hash table to dynamically manage the historical interactions of groups and users. This design supports fundamental operations including insertion, updating, retrieval, forgetting of irrelevant records, and summarization of evolving group and user profiles for efficiently tracking. Based on these retrieved dynamic profiles, the Reason?ing Module then performs a multi-step reasoning process includ?ing Group Interests Collection, Group Consensus Refinement, Multi-dimensional Evaluation and Explainable Recommendation Generation, thereby moving beyond black-box inference to de?liver fully interpretable recommendations. In practice, we adopt the Reinforcement Fine-Tuning (RFT) paradigm, where we first use Supervised Fine-Tuning (SFT) to equip the model with basic capabilities for invoking the Memory and Reasoning modules, and then employ Group Relative Policy Optimization (GRPO) to enhance its autonomous ability to coordinate these modules. Experiments on LastFM and Douban datasets demonstrate that AGR significantly outperforms existing state-of-the-art methods in both recommendation accuracy and explainability. Our model is open-sourced at https://huggingface.co/niuqimeng/AGR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑