发表机构
University of California, Santa Barbara; Microsoft; University of Notre Dame(加州大学圣塔芭芭拉分校; 微软; 圣母大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体长任务中的记忆膨胀问题,提出ReCAP方法,利用持久上下文图存储注意力重要性及依赖关系,无需额外模型调用即可选择相关历史,显著降低压缩延迟并提升代码任务准确率。
AI 中文摘要
随着LLM能力的提升,智能体正在处理越来越复杂、时间跨度更长的任务。它们不断增长的交互相册使得记忆压缩对于保持在上下文窗口内并降低预填充成本至关重要。现有方法总结历史或压缩其KV缓存,通常增加模型计算以保留信息供未来请求使用。新的用户请求可能改变哪些历史是重要的,但如果KV缓存已过期,使用模型重新评估该历史需要重新编码。过去的注意力提供了历史重要性和消息之间依赖关系的信号,而与当前任务的相关性必须使用新的用户请求来评估。我们引入了ReCAP,一种记忆压缩方法,将注意力派生的重要性分数和依赖链接存储在轻量级的持久上下文图中。对于每个新请求,ReCAP将存储的重要性与来自请求的相关性线索相结合,并沿着依赖链接选择消息及其支持上下文,而无需额外的模型调用来进行选择。与Codex默认的基于摘要的压缩相比,ReCAP在Qwen3-Coder和gpt-oss上将压缩和冷恢复的估计延迟降低了约95%。在SWE-Together上,它还将每次调用的历史上下文大约减半,同时保持相当的任务质量,并在Lost-in-Conversation的代码任务上将准确率比完整历史提高了19.8和41.2个百分点。
英文摘要
As LLM capabilities advance, agents are tackling increasingly complex tasks over longer horizons. Their growing interaction histories make memory compaction essential for staying within context windows and reducing prefill cost. Existing methods summarize the history or compress its KV cache, often adding model computation to preserve information for future requests. A new user request can change which history matters, but reassessing that history with the model requires re-encoding it if the KV cache has expired. Past attention provides signals of historical importance and dependencies between messages, while relevance to the current task must be assessed using the new user request. We introduce ReCAP, a memory compaction method that stores attention-derived importance scores and dependency links in a lightweight, persistent context graph. For each new request, ReCAP combines stored importance with relevance cues from the request and follows dependency links to select messages and their supporting context, without additional model calls for selection. Compared with Codex's default summarization-based compaction, ReCAP reduces estimated latency for compaction and cold restoration by approximately 95% on both Qwen3-Coder and gpt-oss. It also roughly halves the historical context per call on SWE-Together at comparable task quality and improves accuracy on the code tasks of Lost-in-Conversation over full history by 19.8 and 41.2 points.