发表机构
Singapore Management University; Nanyang Technological University; Harvard University(新加坡管理大学; 南洋理工大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AutoCompact通过策略学习压缩决策,利用修正轨迹训练并强化学习优化,在SWE基准上显著提升编码智能体通过率。
AI 中文摘要
编码智能体通过代码检查、搜索、编辑和测试的长轨迹来解决仓库级软件工程任务。随着任务进展,早期的探索变得过时,因此管理上下文不仅仅是避免溢出:智能体必须决定何时压缩、保留哪些工作状态以及如何从中继续。我们引入了AutoCompact,它训练编码智能体将这些决策作为其策略的一部分。为了收集训练数据,我们在编码任务上运行基础智能体,并使用评判器审查其压缩决策、摘要和压缩后的操作。有缺陷的输出在环境中执行前会被替换为修正后的版本,因此每条轨迹都从修正后的决策继续。我们使用这些轨迹进行监督微调,然后通过带有任务成功奖励的强化学习联合优化编码和压缩。在SWE-bench Verified和SWE-PolyBench Verified上的实验表明,AutoCompact将基础模型的通过率分别绝对提高了9.2%和5.0%。这些改进在所有评估的推理预算中都保持一致,包括从未溢出的256K上下文窗口和溢出触发回退压缩的16K窗口。
英文摘要
Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2\% and 5.0\%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.