arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

带有快速写入路由与慢速整合的双层智能体记忆

Dual-Layer Agentic Memory with Fast Write Routing and Slow Consolidation

Wenzhi Li, Dong Nie, Rui Lan, Tongtong Lyu, Peiyao Wang, Lingzi Hong, Weihang Pan, Binbin Lin, Boyuan Pan, Yao Hu

arXiv 2608.22215首次发表:更新:

发表机构

Zhejiang University; Xiaohongshu; University of North Texas(浙江大学; 小红书; 北得克萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM智能体动态环境下的记忆管理问题,提出双层智能体记忆框架,通过成本感知路由与周期性整合实现高效记忆处理,在保留高检索性能的同时减少冗余。

AI 中文摘要

大语言模型(LLM)智能体在知识持续演变的动态环境中运行。现有记忆系统通常将外部记忆视为单调增长的存储库,不可避免地会随时间出现检索性能下降和计算成本增加的问题。我们认为核心挑战并非仅在于检索,而在于管理知识生命周期:决定将哪些内容外化、更新或最终内化。受神经科学中的互补学习系统(CLS)理论启发,我们提出双层智能体记忆(Dual-Layer Agentic Memory)框架,通过成本感知的认知路由和周期性参数整合,将记忆管理转移至写入阶段。传入信息被分类为非写入、新写入或更新写入,并通过从小到大的模型级联进行路由,以最小化路由开销同时过滤冗余记忆。后续的写回阶段通过监督微调选择性地将高价值外部记忆整合到模型参数中。实验证明了我们方法的双重效率:17亿参数/80亿参数的级联可修剪多达68%的冗余外部记忆,同时仅提升不到50%的输入,却保留了详尽保留基线所实现的下游问答精确匹配(EM)的98%以上。我们进一步表明,周期性整合可成功内化外部知识,使路由机能在模型认知边界演变时自适应抑制冗余写入。总体而言,我们的框架为智能体记忆提供了统一范式:选择性外化后接选择性内化。代码和数据集将在接收后发布。

英文摘要

Large language model (LLM) agents operate in dynamic environments where knowledge continuously evolves. Existing memory systems typically treat external memory as a monotonically growing repository, inevitably leading to retrieval degradation and increasing computational costs over time. We argue that the core challenge is not retrieval alone, but managing the knowledge lifecycle: deciding what to externalize, update, or ultimately internalize. Inspired by Complementary Learning Systems (CLS) theory in neuroscience, we propose Dual-Layer Agentic Memory, a framework that shifts memory management to the write phase through cost-aware epistemic routing and periodic parametric consolidation. Incoming information is categorized as non-write, write-new, or write-update, and routed through a small-to-large model cascade that minimizes routing overhead while filtering redundant memories. A subsequent write-back phase selectively consolidates high-value external memories into model parameters via supervised fine-tuning. Experiments demonstrate the dual efficiency of our approach: a 1.7B/8B cascade prunes up to 68% of redundant external memory while escalating fewer than 50% of inputs, yet retains over 98% of the downstream QA Exact Match (EM) achieved by an exhaustive retention baseline. We further show that periodic consolidation successfully internalizes external knowledge, allowing the router to adaptively suppress redundant writes as the model's epistemic boundaries evolve. Overall, our framework presents a unified paradigm for agent memory: selective externalization followed by selective internalization. Code and dataset will be released upon acceptance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑