arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32961cs.LGcs.AI

超越令牌节省:LLM智能体中上下文压缩的系统性研究

Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents

Ritul Satish, Prasoon Sinha, Akiho Kawada, Neeraja J. Yadwadkar

首次发表
浏览论文内容

中文总结 AI 辅助

本研究系统解耦LLM智能体上下文压缩的决策维度,通过近35,000次运行发现更少令牌未必更快更省,并强调按任务、模型定制压缩策略。

中文摘要 AI 辅助

随着LLM智能体处理更长的任务,它们越来越多地压缩不断增长的推理、行动和工具输出历史。压缩可以减少令牌使用,但也改变了后续决策可用的信息。现有的智能体框架将压缩什么、何时压缩以及压缩多少的决策捆绑到固定策略中。需要进行系统性表征来理清这些决策,并揭示每个决策如何影响任务成功率和执行成本。我们在SWE-bench Verified和Terminal-Bench 1.0上,对三个开放权重模型系统地变化这些决策。在近35,000次智能体运行中,我们测量了任务成功率、令牌使用量、端到端延迟和估计成本。我们发现,更少的令牌并不一定意味着更快或更便宜的执行:在Terminal-Bench上使用Qwen时,使用约三分之一令牌的策略可能比未压缩的智能体耗时多20-80%。总体成功率相似的策略可以解决不同的任务,而相同的策略在不同模型上的表现可能差异很大。我们的结果促使根据压缩对智能体执行的影响来评估压缩,并根据任务、模型和工作负载定制策略。

英文摘要

As LLM agents tackle longer tasks, they increasingly compress growing histories of reasoning, actions, and tool outputs. Compression can reduce token use, but it also changes the information available for later decisions. Existing agentic harnesses bundle decisions about what to compress, when to compress, and how much to remove into fixed policies. A systematic characterization is needed to disentangle these decisions and reveal how each affects task success and execution cost. We systematically vary these decisions across three open-weight models on SWE-bench Verified and Terminal-Bench 1.0. Across nearly 35,000 agent runs, we measure task success, token use, end-to-end latency, and estimated cost. We find that fewer tokens need not mean faster or cheaper execution: on Terminal-Bench with Qwen, policies using roughly one-third as many tokens can take 20-80% longer than the uncompressed agent. Policies with similar overall success can solve different tasks, while the same policy can perform quite differently across models. Our results motivate evaluating compression by its effects on agent execution and tailoring policies to the task, model, and workload.

↑