arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21503cs.AIcs.IR

智能体上下文管理:将智能体内存和成本视为生命周期和架构问题来解决

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

发表机构Maximem
查看机构详情
  • Maximem

机构由 AI 辅助整理,请以论文原文为准。

Gaurav Dadhich

首次发表
浏览论文内容

中文总结 AI 辅助

研究生产型人工智能智能体因无法管理推理上下文而失败的问题,提出智能体上下文管理(ACM),分解为五个原语,经经济分析和参考实现验证,在特定配置下取得较好结果,还指出了现有基准未涵盖的维度及上下文前沿。

中文摘要 AI 辅助

生产型人工智能智能体的失败往往不是因为推理能力不足,而是因为它们无法管理推理上下文中的内容,如对话历史、大型提示、大型工具定义和不断膨胀的工具输出。现有方法将此视为存储和检索问题,而本文认为这种框架过于狭窄。积极管理智能体的记忆是一个生命周期问题,而不仅仅是存储问题。本文提出智能体上下文管理(ACM),将其分解为五个原语:架构设计、摄取、范围界定、预测以及压缩与整合。通过经济分析表明,只有经过验证的压缩才能在保持保真度的同时实现线性成本。文中描述了一个参考实现Maximem Synap,并报告了在特定配置下LongMemEval达到92%、LoCoMo达到93.2%的结果。最后指出了现有基准尚未涵盖的维度,如延迟、令牌效率和上下文旋转抗性,以及该类别所指向的决策级和组织级上下文的前沿。

英文摘要

Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations. The incumbent response treats this as a storage-and-retrieval problem. We argue that framing is too narrow. Actively managing what an agent holds in mind is a lifecycle, not merely a store: it spans deciding what to remember, extracting and structuring it, choosing the right store per data type, consolidating and forgetting while preserving provenance, deciding what is relevant now, anticipating what is needed next, and compacting context to a budget without losing what matters. In serious production this operates not over a single user but across an organizational scope hierarchy. We name this discipline Agentic Context Management (ACM) and decompose it into five primitives: architecting, ingesting, scoping, anticipating, and compacting & consolidation. We then make the economic case: naive context accumulation grows token cost quadratically in conversation length, crude summarization buys linear cost at the price of an accuracy cliff, and only validated compaction achieves linear cost with preserved fidelity. We describe a reference implementation, Maximem Synap, that realizes the five primitives as a multi-tenant service and reports 92% on LongMemEval and 93.2% on LoCoMo under the configuration detailed in Section 6. We close with dimensions existing benchmarks do not yet capture, latency, token efficiency, and context-rot resistance, and the frontier of decision-level and organization-level context the category points toward.

补充信息

↑