发表机构
University of Miami; Harvard University(迈阿密大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出连续上下文管理(CCM),在每轮压缩交互历史以减少活动提示,并通过特权完整历史蒸馏的强化学习提升性能,验证了其作为低上下文智能体推理范式的可行性。
AI 中文摘要
长视野大型语言模型(LLM)智能体通常在达到预定义阈值时才触发压缩,从而保留其完整的交互历史。我们研究了连续上下文管理(CCM),该方法在每一轮都执行压缩,以防止交互历史在活动提示中累积。在每一轮中,CCM智能体会发出更新的记忆以及环境动作;其下一个提示包含原始任务、保留的记忆和最新观察,而非完整记录。我们首先在TerminalBench-2上使用Claude Sonnet 4.6、Claude Opus 4.6、GLM-5和Kimi K3评估了未经微调的CCM。CCM大幅减少了累积输入使用量和活动提示大小,尽管它降低了对大多数模型的任务成功率,但对Kimi K3保持了性能。我们使用带有特权完整历史蒸馏的GRPO来改进开放权重模型中的CCM。学生初始模型的冻结副本根据该学生轨迹重建的完整历史对每个采样的学生动作进行评分,从而在没有单独教师轨迹或参考解决方案的情况下提供密集的动作令牌监督。在WebShop上,该目标在两种评估模型规模下均显著优于GRPO,并在Qwen3-4B-Instruct上超越了完整历史GRPO,但在Qwen3-8B上未超越。在Endless Terminals上,增强方法相比GRPO提供了适度改进,两种CCM策略均优于未训练的完整历史基线。这些结果表明,CCM对于在显著减少的保留上下文下运行的智能体是一种可行的推理范式,并且其性能可以通过带有特权完整历史蒸馏的强化学习来提高。
英文摘要
Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the active prompt. At each turn, a CCM agent emits an updated memory together with an environment action; its next prompt contains the original task, retained memory, and newest observation rather than the complete transcript. We first evaluate CCM without fine-tuning on TerminalBench-2 using Claude Sonnet 4.6, Claude Opus 4.6, GLM-5, and Kimi K3. CCM substantially reduces cumulative input usage and active-prompt size, although it lowers task success for most models while preserving performance for Kimi K3. We use GRPO with privileged full-history distillation to improve CCM in open-weight models. A frozen copy of the student's initial model scores each sampled student action under the complete history reconstructed from that student's rollout, providing dense action-token supervision without a separate teacher rollout or reference solution. On WebShop, this objective substantially improves CCM over GRPO at both evaluated model scales and surpasses full-history GRPO for Qwen3-4B-Instruct, though not for Qwen3-8B. On Endless Terminals, the augmented method provides a modest improvement over GRPO, with both CCM policies outperforming the untrained full-history baseline. These results demonstrate that CCM is a viable inference paradigm for agents operating with substantially reduced retained context and that its performance can be improved through reinforcement learning with privileged full-history distillation.