Mem++:面向长期组织型LLM智能体的非破坏性记忆
Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents
浏览论文内容
中文总结 AI 辅助
Mem++提出非破坏性记忆框架,通过读取时选择而非写入时压缩,保留完整文档及日期作者,在组织基准上超越现有记忆系统,显著提升长期问答准确性。
中文摘要 AI 辅助
大型语言模型(LLM)智能体现已参与组织工作,其中许多作者在数月内跨文档记录决策。由于修订后的决策以新文档而非编辑的形式出现,回答问题需要知道在给定时间哪个版本有效。然而,大多数记忆系统在写入时压缩记录。通过将每个文档提炼为事实、笔记或图边,这些方法在提出任何问题之前就固定了可回答的内容。为解决这一问题,我们提出Mem++,一种非破坏性记忆框架,将写入时提炼转变为读取时选择。Mem++将每个文档连同其日期和作者完整存储,并且在写入时不调用任何生成模型。在读取时,它仅检索日期截至问题所问时间的文档,并融合词汇和语义排名。与覆盖旧版本的系统不同,Mem++保留旧版本并将选择权留给回答模型。在组织基准OrgMemBench上的评估表明,Mem++在两个回答模型上比最强的记忆系统基线高出8.0至13.1个百分点。使用gpt-4.1-mini时,它还取得了最佳总体得分,比RAG高出2.6分。此外,Mem++在LoCoMo上取得了最佳平均LLM评判得分,并在LongMemEval-S上排名第二,仅次于其实体图变体。基准评估代码可在该https URL获取。
英文摘要
Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at https://github.com/AIDAChip-Inc/mem-plus-plus.
发表机构
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
- AIDAChip Inc.(AIDAChip 公司)
机构由 AI 辅助整理,请以论文原文为准。