arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.05378cs.LG

CompactionRL:面向长视界智能体的带上下文压缩强化学习

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents

Yujiang Li, Zhenyu Hou, Yi Jing, Jie Tang, Yuxiao Dong

首次发表
浏览论文内容

中文总结 AI 辅助

针对长视界大语言模型智能体受上下文窗口限制的问题,提出结合上下文压缩的强化学习方法,联合优化任务执行与摘要生成,在智能体编码任务上取得显著性能提升。

中文摘要 AI 辅助

长视界大语言模型智能体正日益受到有限上下文窗口的限制,因为扩展的交互轨迹可能在任务完成前就超出最大上下文长度。上下文压缩通过汇总先前的交互状态并在压缩后的上下文下继续推演提供了一种自然解决方案,但将压缩整合入强化学习的相关研究仍有待探索。本文提出CompactionRL,一种用于训练具备上下文压缩能力的长视界大语言模型智能体的强化学习策略。该方法通过令牌级损失归一化和跨轨迹广义优势估计,联合优化任务执行与摘要生成。该设计使大语言模型智能体能够从压缩后的长视界轨迹中学习。研究团队在开源模型之上训练CompactionRL,并在智能体编码任务上观测到一致的性能提升。CompactionRL使开源GLM-4.5-Air模型(106B-A30B)在SWE-bench Verified上实现66.8%的Pass@1得分,在Terminal-Bench 2.0上实现24.5%的Pass@1得分,绝对提升幅度分别达到7.0和3.1个百分点。基于GLM-4.7-Flash(30B-A3B)构建的CompactionRL将Pass@1得分分别提升5.5和6.8个百分点,在SWE-bench Verified上达到56.0%,在Terminal-Bench 2.0上达到20.2%。CompactionRL因此已被部署至用于训练开源GLM-5.2模型(750B-A40B)的强化学习流水线中。

英文摘要

Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction offers a natural solution by summarizing previous interaction states and continuing the rollout under a compressed context, but incorporating compaction into reinforcement learning remains underexplored. We propose CompactionRL, a reinforcement learning strategy to train long-horizon agentic LLMs with context compaction. Our approach jointly optimizes task execution and summary generation with token-level loss normalization and cross-segment generalized advantage estimation. This design enables the LLM agents to learn from compacted long-horizon trajectories. We train CompactionRL on top of open models and observe consistent performance gains on agentic coding tasks. CompactionRL enables the open GLM-4.5-Air model (106B-A12B) to achieve Pass@1 scores of 66.4% on SWE-bench Verified and 26.2% on Terminal-Bench 2.0, exceeding the base model under inference-time compaction by 6.6 and 4.9 points, respectively. Built upon GLM-4.7-Flash (30B-A3B), CompactionRL improves Pass@1 by 5.5 and 6.7 points against the base model, reaching 56.0% on SWE-bench Verified and 20.2% on Terminal-Bench 2.0. CompactionRL is thus deployed in the RL pipeline for training the open GLM-5.2 model (750B-A40B).

发表机构

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑