arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在被询问之前记住:MemDream 用于自我探测的记忆演化

Remember Before You're Asked: MemDream for Self-Probing Memory Evolution

Mingfei Lu, Mengjia Wu, Runsong Jia, Zhe Luo, Yi Zhang

arXiv 2609.34545首次发表:更新:

发表机构

Australian Artificial Intelligence Institute (AAII); University of Technology Sydney(澳大利亚人工智能研究院; 悉尼科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有记忆系统缺乏主动测试与修复的问题,提出 MemDream 框架,通过离线梦境周期中三个智能体协同探测、诊断和修复记忆图,并利用策略优化与软衰减机制,在 LoCoMo 和 MAB 基准上显著提升性能。

AI 中文摘要

记忆对于使基于大语言模型的智能体在长时程交互中维持连贯、个性化的行为至关重要。然而,现有的记忆系统存在一个根本性局限:它们从不主动测试自身记忆,仅在真实查询暴露缺陷后才进行修复。这种反应式范式意味着每次检索失败都对应一次已付出代价的真实交互。我们提出 MemDream,一个使大语言模型智能体能够进行自我探测记忆演化的框架。我们的框架定期进入离线梦境周期,在此周期中,三个专门的智能体(Dreamer、Analyst、Consolidator)协同探测、诊断并修复记忆图,在失败发生之前进行干预。通过 Group Relative Policy Optimization 训练的策略学习哪些修复操作能产生持久的检索改进,而软衰减机制则提供由相同预期信号驱动的可逆遗忘。在 LoCoMo 和 MemoryAgentBench 上的实验表明,MemDream 在 LoCoMo 上将答案 F1 提高了 4.5 个点,并在 MAB 上比最强的反应式演化基线获得了高出 9.1 分的总体得分。

英文摘要

Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This reactive paradigm means every retrieval failure corresponds to a real interaction in which the cost has already been paid. We propose MemDream, a framework that enables self-probing memory evolution for LLM agents. Our framework periodically enters offline dream cycles where three specialized agents (Dreamer, Analyst, Consolidator) collaboratively probe, diagnose, and repair the memory graph before failures occur. A policy trained via Group Relative Policy Optimization learns which repair operations produce durable retrieval improvements, while a soft decay mechanism provides reversible forgetting driven by the same anticipatory signal. Experiments on LoCoMo and MemoryAgentBench demonstrate that MemDream improves answer F1 by 4.5 points on LoCoMo and achieves a 9.1-point higher overall score on MAB over the strongest reactive-evolution baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑