DTOC:面向AI智能体自适应上下文管理的动态工具输出压缩
DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents
浏览论文内容
中文总结 AI 辅助
针对上下文窗口限制,提出DTOC框架,通过外部存储完整工具输出并插入可逆占位符,在DeepSWE上降低令牌与步骤,提升解决率并降低成本。
中文摘要 AI 辅助
随着智能体能力的增长,实际限制越来越多地源于受限的上下文窗口,而非模型容量。常见的策略,如截断、启发式老化以及有损摘要,可能会丢弃有用信息或引入幻觉风险。为应对这些挑战,我们提出了动态工具输出压缩(DTOC),一个用于基于LLM的智能体中可扩展上下文管理的框架,该框架将上下文更新建模为智能体推理循环中显式且可逆的操作。DTOC将完整的工具输出保留在外部存储器中,同时在活动上下文中插入紧凑的占位符,从而在需要时实现选择性重建。我们形式化了DTOC机制,将其集成到ReAct风格的智能体架构中,并提供了一个支持按需恢复压缩输出的生产导向实现。在DeepSWE上的实验揭示了模型相关效应:对于响应型模型(Sonnet 4.6、GPT-5.4),DTOC减少了输入令牌(分别减少10.3%和12.7%)和智能体步骤(分别减少2.4%和32.3%),同时提高了解决率(分别提高2.5倍和1.5倍)并降低了每个已解决任务的成本(每个已解决任务的成本分别降低3倍和3.5倍)。对于其他模型,结果更为复杂,GPT-5.5的解决率翻倍且成本减半,但其他模型的解决率无影响且成本有负面影响。消融实验结果表明,可逆性至关重要:仅禁用压缩的变体性能下降,而完整的DTOC在显著降低上下文成本的情况下恢复了基线准确性。这些发现表明,显式、可逆的上下文管理可以在不降低任务性能的情况下提高长程智能体推理的效率。
英文摘要
As agent capabilities have grown, practical limitations increasingly stem from constrained context windows rather than model capacity. Common strategies, such as truncation, heuristic aging, and lossy summarization, may discard useful information or introduce hallucination risk. To address these challenges, we propose Dynamic Tool Output Compression (DTOC), a framework for scalable context management in LLM-based agents that models context updates as explicit and reversible operations within the agent reasoning loop. DTOC retains full tool outputs in external memory while inserting compact placeholders into the active context, enabling selective reconstruction when needed. We formalize the DTOC mechanism, integrate it into a ReAct-style agent architecture, and provide a production-oriented implementation supporting on-demand restoration of compressed outputs. Experiments on DeepSWE reveal model-dependent effects: for responsive models (Sonnet 4.6, GPT-5.4), DTOC reduces input tokens (10.3 and 12.7%) and agent steps (2.4 and 32.3%), while increasing solve rates (2.5 and 1.5 times higher) and lowering cost per solved task (3 and 3.5 times lower cost per solved task). For the other models results are more mixed, with GPT-5.5 doubling solve rate and halving cost, but no impact on solve rate and negative impact on cost for the other models. Ablation results show reversibility is critical: disable-only compression variants degraded performance, while full DTOC recovered baseline accuracy at substantially lower context cost. These findings indicate that explicit, reversible context management can improve the efficiency of long-horizon agent reasoning without degrading task performance.
发表机构
- Pegasystems
- Leiden University(莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。