arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为受控过程的记忆:为大语言模型智能体学习自适应内存管理

Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents

Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu, Xiao Liang, Rui Sun, Yubei Li, Edward Sun, Haozheng Luo, Zhaolu Kang, Aylin Caliskan, Kai-Wei Chang, Ying Nian Wu

arXiv 2607.13591首次发表:更新:

发表机构

University of California Los Angeles; University of Washington; Northwestern University(加利福尼亚大学洛杉矶分校; 华盛顿大学; 西北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究LLM智能体内存管理问题,提出MemCon框架将内存操作建模为马尔可夫决策过程,通过在线策略自适应管理内存,该框架与后端无关,实验表明其在多基准测试中优于基线,提升任务成功率并减少令牌消耗。

AI 中文摘要

大语言模型(LLM)智能体越来越依赖外部内存系统来积累跨任务经验。然而,几乎所有现有方法,从图结构内存到反思洞察存储,都通过固定的、手工设计的启发式方法访问内存。我们认为,这种静态的内存观点是智能体学习的核心瓶颈,因为最佳内存行为本质上依赖于上下文。我们提出了“作为受控过程的记忆”(MemCon)框架,将内存操作建模为马尔可夫决策过程,并学习一种在线策略,以自适应地决定何时、检索什么以及检索多少,何时注入提炼的计划,以及何时进行合并或遗忘。MemCon与后端无关,通过任务级二进制反馈学习,无需预训练和额外的LLM调用。在6个基准测试、3个智能体框架和3个LLM主干上,MemCon在任务成功率上比多个内存基线高出15.2分,同时减少了5%-20%的令牌消耗。

英文摘要

Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of the tasks, benefit from minimal retrieval because memory is sparse; recurring goal types benefit from plan reuse rather than generic nearest-neighbor lookup; stuck agents benefit from re-retrieval with alternative queries; and across long task streams, the memory store itself must be consolidated and pruned to remain useful. We present Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget. MemCon is backend-agnostic: it wraps any existing memory implementation, learns from task-by-task binary feedback with no pretraining and no additional LLM calls, and uses a lightweight tabular contextual bandit with UCB exploration that converges within tens of tasks. Across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MemCon consistently outperforms multiple memory baselines by up to 15.2 points in task success while reducing token consumption by 5--20%.

Commentsnone

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑