arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemHarness:记忆是被重构的,而非被重放的

MemHarness: Memory Is Reconstructed, Not Replayed

Rong Wu, Daocheng Fu, Licheng Wen, Xuemeng Yang, Shu Zou, Jianbiao Mei, Yuxin Wang, Hairong Zhang, Yu Yang, Tao Hu, Cong Zhang, Botian Shi, Pinlong Cai

arXiv 2607.28272首次发表:更新:

AI 中文总结

该研究针对现有记忆增强LLM智能体的逐字重放范式缺陷,提出MemHarness框架,通过GRPO训练实现记忆重构,在ALFWorld、WebShop等任务中性能优于基线,提升了智能体推理与分布外鲁棒性。

AI 中文摘要

检索过往经验已成为增强大型语言模型(LLM)智能体的常用策略。然而,多数现有带记忆增强的智能体将检索到的经验视为需逐字重放的静态记录,无论是否与智能体当前情境匹配,都将其注入上下文。这种“重放”范式忽视了存储经验的抽象通用性与决策时遇到的具体、不断变化的状态之间的差距,常导致负迁移。相比之下,人类很少逐字回忆过往经验,而是重组并调整检索到的记忆以适配当前情境。受此启发,我们提出MemHarness,这一框架使LLM智能体能基于当前情境主动利用和重构过往经验。在每个决策步骤,统一策略模型会基于当前状态评判并重构检索到的经验,在行动前生成基于上下文的指导。这种重构能力通过与GRPO的端到端训练自然显现。在ALFWorld和WebShop上的实验表明,MemHarness的性能显著优于纯强化学习(RL)和静态记忆增强基线,在分布外(OOD)场景中展现出强鲁棒性。此外,我们的分析显示,该重构目标不仅能防止负迁移,还可在训练期间作为潜在指导,从根本上提升智能体的内在推理能力。

英文摘要

Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be replayed verbatim, injecting them into the context regardless of whether they align with the agent's current situation. This ``replay'' paradigm ignores the gap between the abstract, general nature of stored experience and the concrete, ever-changing states encountered at decision time, frequently causing negative transfer. In contrast, humans rarely recall past experiences verbatim; instead, they reorganize and adapt retrieved memories to fit the present context. Inspired by this, we propose MemHarness, a framework that equips LLM agents to actively harness and reconstruct past experiences based on the present context. At each decision step, a unified policy model critiques and reconstructs the retrieved experience conditioned on the current state, producing context-grounded guidance before acting. This reconstructive ability emerges naturally through end-to-end training with GRPO. Experiments on ALFWorld and WebShop show that MemHarness substantially outperforms pure RL and static memory-augmented baselines, demonstrating strong robustness in out-of-distribution (OOD) scenarios. Furthermore, our analyses reveal that this reconstruction objective not only prevents negative transfer but also serves as latent guidance during training, fundamentally improving the agent's intrinsic reasoning capabilities.

Comments20 pages, 13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑