CodeMidas:从代码本身扩展智能体编码强化学习环境
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
另 1 家 · 查看机构详情
- Xiaomi(小米)
- Peking University(北京大学)
- University of Hong Kong(香港大学)
- Renmin University of China(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
CodeMidas提出从源代码自动构建编码RL环境的智能体流水线,生成5,545个跨23种语言的任务,训练后显著提升多种编码基准性能。
中文摘要 AI 辅助
通过强化学习(RL)训练有能力的编码智能体需要多样化的任务和可靠的验证器。开源代码库提供了此类任务的丰富来源,而现有方法通常依赖于问题和提交等开发工件,限制了可提取任务的范围。为了更好地扩展强化学习环境,我们提出了CodeMidas,一个智能体流水线,它将现有代码库中已实现的功能转化为可执行的强化学习环境,仅使用源代码作为其任务特定的输入。CodeMidas将智能体计算分配给环境构建的每个阶段:智能体探索已实现的功能以制定行为规范,构建基于原始代码执行的测试,并通过执行检查和重复的解决方案展开来验证和过滤候选任务。由此产生的数据集包含来自3,185个开源代码库的5,545个训练任务,涵盖23种编程语言和15个技术领域。在这些任务上使用GRPO训练MiMo-V2.5,在五个不同的基准测试上均提升了性能,涵盖问题修复(DeepSWE +11.7%)、整个程序构建(ProgramBench +17%)和终端工作(Terminal-Bench v2.1 +8.5%)。消融实验表明,增加高质量训练任务的数量可提升性能。轨迹分析显示,经过强化学习训练的智能体表现出更好的行为,如增加代码库探索和更多样化的自我验证。这些结果确立了源代码作为构建强化学习环境的可扩展基础,可提升编码智能体在多样化软件任务中的表现。
英文摘要
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input. CodeMidas allocates agentic compute to every stage of environment construction: agents explore implemented functionality to formulate behavioral specifications, construct tests grounded in execution of the original code, and validate and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on these tasks with GRPO improves performance on all five diverse benchmarks, covering issue repair (DeepSWE + 11.7%), whole-program construction (ProgramBench +17%), and terminal work (Terminal-Bench v2.1 +8.5%). Ablations show that increasing the number of high-quality training tasks improves performance. Trajectory analysis shows the RL-trained agent demonstrates better behaviors like increasing codebase exploration and more diverse self-verification. These results establish source code as a scalable foundation for constructing RL environments that improve coding agents across diverse software tasks.