AI 中文总结
研究针对工具使用智能体的控制上下文压缩,提出CompressAgent基准,揭示压缩的非线性可靠性边界,发现压缩会引发工具执行等错误,需以可执行结果评估压缩效果。
AI 中文摘要
使用工具的语言模型智能体不仅受任务提示控制,还受指定工具、参数、策略、执行协议及恢复方式的持续系统侧指令控制。压缩这些智能体控制上下文(ACC)可降低输入成本和上下文使用量,但现有提示压缩评估未揭示压缩后的控制是否仍保持操作可靠性。我们推出CompressAgent,这是一个经环境验证的ACC压缩基准,涵盖9个独立构建的ACC、3个任务族、3个固定Qwen API模型标识符、6个保留上下文预算及15525次运行。我们发现非线性的、依赖方法的可靠性边界:在75%保留上下文时,通用重写和基于分段的压缩分别达到92.7%和92.4%的成功率,接近93.8%的全上下文基线;在50%至35%区间,方法差异显著,35%保留上下文时,基于分段、感知义务及通用重写的成功率分别为47.0%、39.0%和19.9%;在25%至10%的保留上下文预算下,可执行协议变得脆弱。可靠性还随ACC大幅变化,因此通用压缩器排名不合适,需按上下文限定。故障分析显示,压缩主要表现为工具执行和动作解析错误。这些发现将ACC压缩从token减少问题重新定义为必须通过可执行结果评估的运行时可靠性问题。
英文摘要
Tool-using language-model agents are governed not only by task prompts but also by persistent system-side instructions that specify tools, arguments, policies, execution protocols, and recovery. Compressing these agent control contexts (ACCs) can reduce input cost and context use, yet existing prompt-compression evaluations do not reveal whether the resulting control remains operationally reliable. We introduce CompressAgent, an environment-verified benchmark for ACC compression across nine independently constructed ACCs, three task families, three fixed Qwen API model identifiers, six retained-context budgets, and 15,525 runs. We uncover a nonlinear, method-dependent reliability frontier. At 75% retained context, generic rewriting and section-based compression achieve 92.7% and 92.4% success, close to the 93.8% full-context baseline. Between 50% and 35%, methods diverge sharply; at 35%, section-based, obligation-aware, and generic rewriting achieve 47.0%, 39.0%, and 19.9%. At retained-context budgets from 25% to 10%, executable protocols become fragile. Reliability also varies substantially across ACCs, making universal compressor rankings inappropriate and motivating per-context qualification. Failure analysis shows that compression primarily surfaces as tool-execution and action-parsing errors. These findings recast ACC compression from token reduction into a runtime-reliability problem that must be evaluated through executable outcomes.
Comments12 pages, 5 figures; includes an appendix