arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文压缩会给智能体带来什么代价?任务完成指标未揭示的交互代价

What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics

Shuyu Liu

arXiv 2608.16370首次发表:更新:

发表机构

Chongqing University of Science and Technology(重庆科技学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对使用工具的智能体,发现上下文压缩会带来任务完成指标未揭示的隐藏交互成本,检索调用增加是其主要来源,且该特征具有环境依赖性。

AI 中文摘要

任务完成是评估上下文压缩的标准指标,但它并不全面:压缩会迫使智能体重新获取被丢弃的状态,从而增加其交互成本,同时在统计上不改变完成情况。我们为有限 horizon 下使用工具的智能体引入了一种受控的运行时测量协议,该智能体在确定性规划环境中以固定的 24 步 horizon 运行。我们改变压缩程度,比较丢弃算子与保留事实的算子,通过受控的神谕干预恢复被丢弃的状态,并将工具调用分解为检索和执行。我们在两种任务 regime 中评估了三种模型。在所有 6 个模型-regime 比较中,检索调用均增加,且几乎占了所有额外交互的全部;在 Holm 校正后,其中 5 个仍具有统计学意义。在预先指定的 5 倍比较点,任何单元中的完成变化均无统计学意义;DeepSeek 仅在 10 倍压缩时显示出显著的完成下降。GPT-5.5 是最明显的案例:完成率从 80% 变为 85%(p=1.0),而检索调用从 21.0 增加到 63.9 次(p=0.002)。保留干预进一步区分了状态数量、状态类型和内容有效性:随机选择与离线事后神谕相当,而用语义不相关内容替换保留的 D-状态会使检索增加 57%(p<0.001),且无显著的完成变化。在第二个环境 ALFWorld 中,滑动压缩未产生检索激增,表明重新获取特征是环境依赖的,而非缩短上下文的固有属性。总体而言,当与执行相关的状态缺失且必须重新获取时,压缩会带来隐藏的交互成本,而仅靠完成指标可能无法揭示这些成本。

英文摘要

Task completion is the standard metric for evaluating context compression, yet it is incomplete: compression can increase an agent's interaction cost by forcing it to reacquire dropped state while leaving completion statistically unchanged. We introduce a controlled runtime measurement protocol for reacquisition cost in a bounded-horizon tool-using agent. The agent acts in a deterministic planning environment under a fixed 24-turn horizon. We vary compression severity, compare a dropping operator with a fact-preserving operator, restore dropped state through controlled oracle interventions, and decompose tool calls into retrieval and execution. We evaluate three models across two task regimes. Retrieval calls increase in all six model-regime comparisons and account for almost all added interaction; five of six remain significant after Holm correction. At the prespecified 5x comparison point, completion changes are not significant in any cell. DeepSeek shows a significant completion drop only at 10x compression. GPT-5.5 is the clearest case: completion changes from 80% to 85% (p = 1.0) while retrieval increases from 21.0 to 63.9 calls (p = .002). Retention interventions further separate state quantity, state type, and content validity. Random selection is comparable to an offline hindsight oracle, while replacing retained D-state with semantically irrelevant content increases retrieval by 57% (p < .001) without a significant completion change. In a second environment, ALFWorld, sliding compression produces no retrieval surge, showing that the reacquisition signature is environment-dependent rather than intrinsic to shortening context. Overall, compression can impose hidden interaction costs when execution-relevant state becomes absent and must be reacquired, while completion alone may not expose those costs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑