发表机构
Rice University; IBM Research(莱斯大学; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM在线持续学习中的上下文膨胀问题,提出图结构记忆GraphMemory,按查询检索相关子图,在保持性能的同时减少81-85%的记忆构建令牌。
AI 中文摘要
大型语言模型(LLMs)正越来越多地部署于企业、科学和医疗应用中,在这些场景中,智能体必须整合领域特定知识并从经验中适应。上下文工程通过推理时提供的指令、策略和证据来改进模型行为,为权重更新提供了一种实用替代方案。然而,在线调整上下文通常需要昂贵的试错过程,而查询往往被独立处理,阻碍了有用经验的传递。记忆系统通过跨交互保留信息来解决这一限制,但持续将信息追加到共享上下文的方法会面临不断增加的令牌成本、上下文窗口限制以及随着上下文扩展导致的性能下降。我们引入了上下文优化的统一表述,并证明智能体记忆系统的更新可以被解释为对模型上下文的一种优化更新过程。这一视角试图为研究记忆设计及其效率提供一个原则性框架。随后,我们提出了GraphMemory,一种轻量级的基于图的记忆,它能够积累、精炼、组织并连接可复用的策略。对于每个查询,GraphMemory仅检索相关的子图,从而实现在不将整个记忆暴露给模型的情况下进行在线上下文适应。在有界检索条件下,随着已处理示例数量的增长,检索到的记忆量保持恒定。实验表明,GraphMemory在实现具有竞争力的下游性能的同时,相比我们的基线方法,使用的记忆构建令牌大约减少了81-85%。
英文摘要
Large language models (LLMs) are increasingly deployed in enterprise, scientific, and medical applications, where agents must incorporate domain-specific knowledge and adapt from experience. Context engineering offers a practical alternative to weight updates by improving model behavior through instructions, strategies, and evidence supplied at inference time. However, adapting context online typically requires a costly trial-and-error process, while queries are often processed independently, preventing useful experience from carrying forward. Memory systems address this limitation by retaining information across interactions, but approaches that continually append information to a shared context face increasing token costs, context-window limits, and performance degradation as the context expands. We introduce a unified formulation of context optimization and show that an agent memory system update can be interpreted as an optimization update procedure over the model's context. This perspective attempts to provide a principled framework for studying memory design and its efficiency. We then propose GraphMemory, a lightweight graph-based memory that accumulates, refines, organizes, and connects reusable strategies. For each query, GraphMemory retrieves only the relevant subgraph, enabling online context adaptation without exposing the model to the entire memory. Under bounded retrieval, the amount of retrieved memory remains constant as the number of processed examples grows. Experiments show that GraphMemory achieves competitive downstream performance while using approximately 81-85% fewer memory-construction tokens than our baselines.
Comments14, 4, neurips workshop: TTCL