编码智能体记忆后训练:通过强化学习解锁预训练文件操作的长时程任务记忆潜力
Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对长时程任务中智能体记忆不足的问题,提出CAMG环境套件和CAMG-RL训练方法,利用文件操作和强化学习提升记忆能力,在基准测试中达到与更大模型相当的性能。
中文摘要 AI 辅助
语言模型智能体日益频繁地处理长时程任务,其交互历史往往超出模型的主动上下文窗口。近期研究开始利用强化学习将记忆控制纳入策略,但通常依赖于领域特定训练环境中预定义的记忆工具,且这些环境的时间跨度相对较短。这种设置将学习到的记忆行为与基础模型预训练之外的环境特定接口绑定,必须从零开始学习,因此即使在经过后训练后,智能体在长时程任务中仍难以有效运用记忆。为解决这些局限,我们提出了编码智能体记忆健身房(CAMG),一个涵盖购物、编码、深度研究和自动研究的长时程智能体强化学习环境套件。除各环境原生任务接口外,CAMG还提供可执行的shell访问和跨回合持久的工作空间,使智能体能够在整个回合中创建、修改、搜索和重用文件作为记忆。我们还引入了CAMG-RL,它使用完全异步的PPO在全部四个环境上联合训练单一策略,直接从下游任务奖励中学习这种基于文件的记忆行为,并从相应规模的Qwen3.5模型训练出CAMG-RL-4B和CAMG-RL-9B。在SWE-bench Verified和MLE-bench Lite上,CAMG-RL-4B和CAMG-RL-9B分别与Qwen3.5-35B-A3B和Qwen3.5-122B-A10B具有竞争力。
英文摘要
Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons. This setup ties learned memory behavior to environment-specific interfaces that lie outside the base model's pre-training and must be learned from scratch, so even after post-training, agents struggle to use memory in long-horizon tasks. To address these limitations, we introduce Coding Agent Memory Gym (CAMG), a suite of long-horizon agentic-RL environments spanning Shop, Coding, DeepResearch, and AutoResearch. Alongside each environment's native task interface, CAMG provides executable shell access and an episode-persistent workspace, enabling agents to create, revise, search, and reuse files as memory throughout an episode. We also introduce CAMG-RL, which trains a single policy jointly across all four environments with fully asynchronous PPO, learning this file-based memory behavior directly from downstream task reward, and we train CAMG-RL-4B and CAMG-RL-9B from Qwen3.5 models of matching size. On SWE-bench Verified and MLE-bench Lite, CAMG-RL-4B and CAMG-RL-9B are competitive with Qwen3.5-35B-A3B and Qwen3.5-122B-A10B, respectively.