发表机构
University of California Los Angeles; University of Washington; Northwestern University(加利福尼亚大学洛杉矶分校; 华盛顿大学; 西北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究LLM智能体内存管理问题,提出MemCon框架将内存操作建模为马尔可夫决策过程,通过在线策略自适应管理内存,该框架与后端无关,实验表明其在多基准测试中优于基线,提升任务成功率并减少令牌消耗。
AI 中文摘要
大语言模型(LLM)智能体越来越依赖外部内存系统来积累跨任务经验。然而,几乎所有现有方法,从图结构内存到反思洞察存储,都通过固定的、手工设计的启发式方法访问内存。我们认为,这种静态的内存观点是智能体学习的核心瓶颈,因为最佳内存行为本质上依赖于上下文。我们提出了“作为受控过程的记忆”(MemCon)框架,将内存操作建模为马尔可夫决策过程,并学习一种在线策略,以自适应地决定何时、检索什么以及检索多少,何时注入提炼的计划,以及何时进行合并或遗忘。MemCon与后端无关,通过任务级二进制反馈学习,无需预训练和额外的LLM调用。在6个基准测试、3个智能体框架和3个LLM主干上,MemCon在任务成功率上比多个内存基线高出15.2分,同时减少了5%-20%的令牌消耗。
英文摘要
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of the tasks, benefit from minimal retrieval because memory is sparse; recurring goal types benefit from plan reuse rather than generic nearest-neighbor lookup; stuck agents benefit from re-retrieval with alternative queries; and across long task streams, the memory store itself must be consolidated and pruned to remain useful. We present Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget. MemCon is backend-agnostic: it wraps any existing memory implementation, learns from task-by-task binary feedback with no pretraining and no additional LLM calls, and uses a lightweight tabular contextual bandit with UCB exploration that converges within tens of tasks. Across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MemCon consistently outperforms multiple memory baselines by up to 15.2 points in task success while reducing token consumption by 5--20%.
Commentsnone