arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24097cs.AI

MemChain:为内存增强型大语言模型智能体学习可解释的内存痕迹

MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学交叉科学学院)
  • Memorax AI(Memorax人工智能公司)
  • School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

Yiwen Ma, Songjun Tu, Qichao Zhang, Dong Li, Linjing Li, Dongbin Zhao

AI总结:

研究针对内存增强型大语言模型智能体检索内存回答查询存在的问题,提出MemChain可训练检索后内存策略,经两阶段学习框架训练,能生成有效证据上下文,实验证明其在多种模型上性能领先且减少内存上下文。

AI中文摘要:

内存增强型大语言模型智能体通常通过检索相关内存并将其直接输入答案模型来回答查询。这种以检索为证据的范式假设检索到的内存已经适合推理,而让答案模型解决冗余、冲突和弱相关性问题,同时在长期内存任务中产生大量上下文开销。我们提出了MemChain,一种可训练的检索后内存策略,它将检索到的候选内存转换为面向答案的活跃内存,以紧凑且有根据的证据上下文表示。给定用户查询和检索到的候选内存,MemChain首先生成一个基于问题的证据计划,然后构建一个有序的有根据的证据痕迹,根据语义角色和依赖关系组织检索到的内存,最后执行显式内存操作以生成用于答案生成的简洁证据上下文。为了训练调解器,我们引入了一个两阶段学习框架。监督痕迹学习首先教导策略生成结构有效的计划、痕迹、操作和证据上下文。然后我们提出了痕迹引导的内存策略优化(TMPO),这是一种强化学习目标,它使用下游答案质量优化内存策略,同时在多个展开中共同鼓励痕迹基础、证据支持、结构有效性和答案稳定性。在LoCoMo和LongMemEval-S上的实验表明,MemChain在闭源和开放权重冻结答案模型上都始终实现了领先的性能,同时大幅减少了传递给答案模型的内存上下文。

英文摘要:

Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model. This retrieval-as-evidence paradigm assumes retrieved memories are already suitable for reasoning, leaving the answer model to resolve redundancy, conflicts, and weak relevance while incurring substantial context overhead in long-term memory tasks. We propose MemChain, a trainable post-retrieval memory policy that transforms retrieved candidates into answer-facing active memory, represented as a compact and grounded evidence context. Given a user query and retrieved candidates, MemChain first generates a question-conditioned evidence plan, then constructs an ordered grounded evidence trace that organizes retrieved memories according to their semantic roles and dependencies, and finally executes explicit memory actions to produce a concise evidence context for answer generation. To train the mediator, we introduce a two-stage learning framework. Supervised trace learning first teaches the policy to generate structurally valid plans, traces, actions, and evidence contexts. We then propose Trace-Guided Memory Policy Optimization (TMPO), a reinforcement learning objective that optimizes the memory policy using downstream answer quality while jointly encouraging trace grounding, evidence support, structural validity, and answer stability across multiple rollouts. Experiments on LoCoMo and LongMemEval-S demonstrate that MemChain consistently achieves state-of-the-art performance across both closed-source and open-weight frozen answer models while substantially reducing the memory context passed to the answer model.

↑