arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

rEDMRec:将大语言模型推理提炼为可编辑的体验记忆用于推荐

rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation

Minh Hoang Nguyen, Tung Le, Huy Tien Nguyen

arXiv 2608.18952首次发表:更新:

发表机构

Faculty of Information Technology, University of Science; Vietnam National University(胡志明市科学大学信息学院; 越南国家大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

rEDMRec将教师LLM推理提炼为四类可编辑体验记忆,轻量级学生LLM通过检索记忆排序,在多数据集和主干模型上优于零样本、少样本、RAG及GraphRAG,记忆优化可提升推荐性能并降低重复率。

AI 中文摘要

大语言模型可通过对用户历史和候选物品进行显式推理来提升推荐质量,例如提取用户偏好或解释为何某一物品更合适,而非直接将历史映射为排序列表。然而,这种推理在每次排序请求时重复计算成本高昂,且生成后通常仅使用一次就被丢弃,既无法在未来请求中复用,也难以在用户偏好变化时检查或修正。我们的核心见解是,若能将推理一次性压缩为紧凑结构化的记忆,由轻量级模型从中检索,就无需在每次调用时重新生成。我们提出rEDMRec,它将教师大语言模型的推理提炼为四个类型化、可编辑的体验通道——长期偏好、短期上下文、物品感知和反事实难负样本比较,由大语言模型记忆控制器执行添加、删除、修改、保留操作,并通过K智能体辩论优化条目。轻量级学生大语言模型仅通过检索该记忆对候选物品排序,无需再次调用教师模型,将在线推理成本与推理深度解耦。在ML-1M、Amazon Beauty、Steam数据集及十个学生主干模型上,rEDMRec在所有主干模型上均优于零样本、少样本和RAG方法,在多数主干模型上优于GraphRAG方法,在ML-1M上较次优基准的提升幅度最高达13.3%。通道 ablation 实验显示,短期上下文是唯一在所有容量层级均持续有效的通道,而长期、物品感知和反事实贡献则依赖容量(在最强学生模型上甚至可能产生反向效果);基于辩论的记忆优化将记忆库重复率降低7.4个百分点,在六个优化周期内将下游HR@1指标提升最高达0.029。

英文摘要

Large language models can improve recommendation quality by reasoning explicitly over user history and candidate items - for example, extracting a user's preferences or explaining why one item fits better than another - rather than mapping history directly to a ranked list. This reasoning, however, is expensive to repeat on every ranking request and, once produced, is typically consumed once and discarded, leaving it neither reusable across future requests nor easy to inspect or correct as user tastes drift. Our insight is that reasoning does not need to be regenerated at every call if it can instead be compressed once into a compact, structured memory that a lightweight model retrieves from. We propose rEDMRec, which distills a teacher LLM's reasoning into four typed, editable experience channels - long-term preference, short-term context, item-perception, and counterfactual hard-negative comparisons - maintained by an LLM memory controller that performs Add/Delete/Modify/Keep operations and refines entries via K-agent debate. A lightweight student LLM then ranks candidates purely by retrieving from this memory, without invoking the teacher again, decoupling online inference cost from reasoning depth. Across ML-1M, Amazon Beauty, and Steam and ten student backbones, rEDMRec improves HR@1 over zero-shot, few-shot, and RAG on every backbone, and over GraphRAG on most backbones, with Impv up to 13.3% vs. the second-best baseline on ML-1M. Channel ablations show that short-term context is the only channel that helps consistently across capacity tiers, whereas long-term, item-perception, and counterfactual contributions are capacity-dependent (and can reverse on the strongest students); debate-based memory optimization lowers bank duplication by 7.4 percentage points while raising downstream HR@1 by up to +0.029 over six optimization epochs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑