发表机构
Surge AI(Surge AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过专家构建的智能体编码任务进行强化学习后训练Kimi K2.7 Code模型,显著提升了六个外部基准的pass@1,并验证了所学技能可跨基准迁移,同时缩短了轨迹并避免了常见失败模式。
AI 中文摘要
编码智能体常常在最后一步失败:它们构建了大部分功能却遗漏某个需求,只测试其实现已经能处理的用例,破坏本应保持不变的行为,或基于未经检验的假设进行验证。我们研究在专家构建的智能体编码任务上进行强化学习(RL)是否能弥合这一差距,以及智能体所学到的知识是否能迁移到训练分布之外。我们使用仅强化学习的方法,在1,700个任务上对Kimi K2.7 Code(一个1T参数(32B激活)的开源混合专家模型)进行后训练:1,000个仓库任务通过隐藏的失败到通过(fail-to-pass)测试和现有行为的通过到通过(pass-to-pass)测试进行评分,以及700个终端任务通过专家编写的隐藏验证器进行评分。奖励是目标检查通过的比例,如果任何通过到通过测试失败,奖励将降为零。在秩为32的LoRA适配器上进行一轮GSPO训练,使我们在三个智能体框架下评估的六个外部基准上的pass@1均有所提升:SWE-Bench Pro(60.1提升至64.8)、DeepSWE(31.0提升至43.4)、Terminal-Bench 2.1(67.4提升至82.0)、Terminal-Bench 3(1.4提升至12.1)、Terminal-Bench 4(0.0提升至7.6)和SWE-Marathon(5.0提升至25.0)。合并五个独立任务集(Terminal-Bench 4修订了Terminal-Bench 3)后,改进显著(p < 0.001),并且在训练数据收集后发布的三个任务集上仍然显著(p = 0.004);该模型在训练中从未使用的两个框架下也有所改进。DeepSWE和Terminal-Bench 3上的中位轨迹在智能体步骤上缩短了24-35%。基础模型失败的DeepSWE运行大多是接近成功的,而在训练模型新解决的任务上,成对轨迹显示它避免了上述四种失败模式。
英文摘要
Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unchecked assumption. We ask whether reinforcement learning (RL) on expert-built agentic coding tasks closes this gap, and whether what the agent learns transfers beyond the training distribution. We post-train Kimi K2.7 Code, a 1T-parameter (32B active) open-weight mixture-of-experts model, with RL alone on 1,700 tasks: 1,000 repository tasks graded by hidden fail-to-pass tests and by pass-to-pass tests of existing behavior, and 700 terminal tasks graded by expert-written hidden verifiers. The reward is the fraction of target checks passed and drops to zero if any pass-to-pass test fails. One epoch of GSPO on a rank-32 LoRA adapter improves pass@1 on each of the six external benchmarks we evaluated, across three agent harnesses: SWE-Bench Pro (60.1 to 64.8), DeepSWE (31.0 to 43.4), Terminal-Bench 2.1 (67.4 to 82.0), Terminal-Bench 3 (1.4 to 12.1), Terminal-Bench 4 (0.0 to 7.6), and SWE-Marathon (5.0 to 25.0). Pooled over the five independent task sets (Terminal-Bench 4 revises Terminal-Bench 3), the improvement is significant (p < 0.001), and it remains significant on the three sets released after the training data was collected (p = 0.004); the model also improves under both harnesses never used in training. Median trajectories on DeepSWE and Terminal-Bench 3 are 24-35% shorter in agent steps. The base model's failed DeepSWE runs are mostly near-misses, and on the tasks the trained model newly solves, paired trajectories show it avoiding each of the four failure modes above.
Comments15 pages, 2 figures, 4 tables