发表机构
Salesforce AI Research(赛富时人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究代理历史能否预测上下文压缩的损害,发现预测能力弱,最佳触发器避免21%有害边界并保留84%压缩机会,但需更多语料库验证。
AI 中文摘要
许多长视界代理基于全局规则(通常是令牌预算)压缩其上下文,而对代理正在执行的任务视而不见。我们探究代理的近期行为是否能预示压缩何时会造成损害。TRACE公开的语料库包含590个由测试平台触发的AppWorld压缩边界,每个边界在压缩前上下文和摘要下从重新执行的前缀状态进行重放,并记录后续动作的负担:出错或重复已执行调用的调用。我们发现边界前的历史仅能弱预测压缩后的损害。一个内部预设的按前缀位置对比是广泛的零结果,其背后的朴素“已写入”标签实际上衡量的是轨迹阶段。最佳扩展协议触发器达到保留测试AUROC 0.66(在复制品自身标签上为0.64),而同一边界的复制品为0.72;最佳冻结可解释触发器避免了21%的有害(正负担)边界,同时保留了84%的压缩机会,并在计数上超过随机规则期望,但在负担质量上未超过(事后比较)。最佳触发器在匹配保留率下是否优于令牌预算规则,无法在发布版本上评估。我们说明了语料库应包含哪些内容以回答该问题。
英文摘要
Many long-horizon agents compact their context on a global rule, usually a token budget, blind to what the agent was doing. We ask whether the agent's recent behaviour predicts when a compaction will hurt. TRACE's public corpus of 590 harness-triggered AppWorld compaction boundaries replays each boundary from a re-executed prefix state under the pre-compaction context and under the summary, and records the burden of the next actions: calls that error or repeat a call already made. We find that pre-boundary history predicts post-compaction harm only weakly. An internally prespecified contrast by prefix placement is a wide null, and the naive "has-written" label behind it turns out to measure trajectory phase. The best extension-protocol trigger reaches held-out AUROC 0.66 (0.64 on the replicate's own label) against a same-boundary replicate of 0.72; the best frozen, interpretable trigger avoids 21% of harmful (positive-burden) boundaries while keeping 84% of compaction opportunities, and exceeds the random-rule expectation on count but not on burden mass (a post hoc comparison). Whether the best trigger beats a token-budget rule at matched retention cannot be evaluated on the release. We state what corpora should ship to answer it.
CommentsAccepted at the IAB Workshop (Interpreting Agent Behavior) at NeurIPS 2026 (non-archival). 20 pages