arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27298cs.AI

StateComp:学习在长时程智能体中何时压缩历史

StateComp: Learning When to Compress History in Long Horizon Agents

Mingxuan Wang, Hongyue Chen, Yinglong Guo, Fei Luo, Chao Ning, Bo Wang, Guorun Yao, Yanbiao Ma, Jungong Han

首次发表
浏览论文内容

中文总结 AI 辅助

StateComp提出状态条件压缩框架,根据智能体当前状态决定历史交互的压缩时机,在WorkBuddyBench上减少52.27%令牌并保持性能,实现12.67倍加速。

中文摘要 AI 辅助

长时程智能体在任务执行过程中会持续积累交互历史,然而随着智能体状态的演变,过往交互的重要性也在不断变化。现有的上下文管理方法大多基于固定窗口、周期性调度或当前相关性来压缩历史,却忽略了一个更根本的问题:何时过去的交互已变得可以安全替换?过早压缩可能会移除未来行动仍所需的信息,而过于保守的保留则会导致大量的上下文开销。为解决这一问题,我们提出了状态条件压缩(StateComp),一个根据当前智能体状态决定历史交互何时可以被安全压缩的框架。StateComp通过两阶段标注流程构建KEEP和READY监督信号,并在冻结语言模型的隐藏表示上训练一个对不平衡敏感的路由器。一个有界的状态表示进一步降低了评估长历史的成本,同时相邻的READY交互被分组为连续片段,并在执行过程中替换为紧凑摘要。在WorkBuddyBench上的实验表明,StateComp在保持任务性能的同时,将智能体和摘要的总令牌数减少了52.27%,并在表示提取上实现了12.67倍的加速。

英文摘要

Long-horizon agents continuously accumulate interaction history during task execution, yet the importance of past interactions changes as the agent state evolves. Existing context management methods largely compress history based on fixed windows, periodic schedules, or current relevance, overlooking a more fundamental question: when has a past interaction become safe to replace? Premature compression may remove information still needed for future actions, while overly conservative retention leads to substantial context overhead. To address this, we propose State Conditioned Compression (StateComp), a framework that determines when historical interactions can be safely compressed according to the current agent state. StateComp constructs KEEP and READY supervision through a two-stage annotation procedure and trains an imbalance-aware router on hidden representations from a frozen language model. A bounded state representation further reduces the cost of evaluating long histories, while adjacent READY interactions are grouped into continuous spans and replaced with compact summaries during execution. Experiments on WorkBuddyBench show that StateComp reduces total agent and summarization tokens by 52.27% while maintaining task performance, and achieves a 12.67-fold speedup in representation extraction.

发表机构

  • TierFlow Team(TierFlow团队)
  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑