AI 中文总结
研究针对长时程语言模型代理上下文窗口超界问题,提出ARC框架,将存档与活动上下文分离,以ID可寻址日志存储工具观察结果,压缩时替换旧观察结果。实验表明该方法能提高信息保留和服务效率。
AI 中文摘要
长时程语言模型代理会积累推理痕迹、动作和工具观察结果,最终可能超出模型的固定上下文窗口。现有压缩方法存在丢弃关键细节或无法可靠恢复的问题。我们提出了ARC(可寻址召回压缩)框架,将存档存储与活动上下文呈现分离。ARC将工具观察结果存储在只追加、ID可寻址的日志中,压缩时用紧凑引用替换旧观察结果。通过Qwen3-8B和Qwen3-32B评估,在“大海捞针”评估中,ARC平均精确答案准确率达99.40%,在LongBench-v2 Hard子集上平均准确率为29.97%。结果表明基于地址的召回可提高信息保留和服务效率。
英文摘要
Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window. Existing compaction methods address this limitation by discarding, summarizing, or retrieving earlier information, but they may remove task-critical details or fail to recover them reliably. We propose ARC (Addressable Recall Compaction), a context-management framework that separates archival storage from active-context presentation. ARC stores tool observations in an append-only, ID-addressable log and replaces older observations with compact citations when compaction is required. The agent can subsequently use these identifiers to request stored content without re-executing the corresponding tools or depending solely on similarity-based retrieval. We evaluate ARC using Qwen3-8B with a 16k context window and Qwen3-32B with a 32k context window. On the Needle-in-a-Haystack evaluation, ARC achieves an average exact-answer accuracy of 99.40%, compared with 88.12% for the best-performing baseline in our evaluation. ARC also reduces estimated serving time and HBM traffic under our hardware-cost model. On the LongBench-v2 Hard subset, ARC obtains an average accuracy of 29.97%, compared with 28.25% for the best-performing baseline. These results indicate that explicit, address-based recall can improve information retention and serving efficiency relative to the evaluated context-management baselines under the tested settings.
Comments20 pages, 2 figures