arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向长 horizon 智能体的可靠上下文压缩:执行不稳定性的实证研究

Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability

Guanghui Min, Liang Wu, Mayank Darbari, Chen Chen, Liangjie Hong

arXiv 2608.06503首次发表:更新:

发表机构

University of Virginia; Nokia(弗吉尼亚大学; 诺基亚公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对长 horizon 智能体上下文压缩的执行不稳定性问题,提出了验证器引导框架 TRACE,在 AppWorld 上的实验显示其在任务性能等指标上优于现有基线,为可靠上下文压缩提供了新方向。

AI 中文摘要

循环上下文压缩控制长 horizon 智能体的上下文增长,但其行为影响仍未被充分理解。在这项初步实证研究中,我们表明压缩会削弱近期交互的影响,增加被阻止的动作、重复探索以及不同运行间的不稳定性。基于这些观察,我们提出了 TRACE,一种验证器引导的框架,它通过来自同一环境状态的成对闭环延续来评估单个压缩事件,并使用摘要偏好来优化自然语言压缩提示,同时保持所有模型冻结。在 AppWorld 上的初步结果显示,与现有压缩基线相比,TRACE 在任务性能、多运行可靠性以及上下文-执行效率方面均有所提升。这些发现为边界局部评估作为可靠智能体上下文压缩的有前景方向提供了早期证据。

英文摘要

Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood. In this preliminary empirical study, we show that compression can weaken the influence of recent interactions, increasing blocked actions, repeated exploration, and instability across runs. Motivated by these observations, we introduce TRACE, a verifier-guided framework that evaluates individual compaction events through paired closed-loop continuations from the same environment state and uses summary preferences to optimize a natural-language compression prompt while keeping all models frozen. Initial results on AppWorld show improvements over existing compression baselines in task performance, multi-run reliability, and context--execution efficiency. These findings provide early evidence for boundary-local evaluation as a promising direction for reliable agent context compression.

Comments31 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑