分组智能体评分与优势再分配用于代码智能体强化学习
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
AI总结:
提出GAGAR框架,通过智能体评分器对组内通过测试的轨迹进行排名并重新分配优势,提升代码智能体强化学习的质量与稳定性。
AI中文摘要:
强化学习(RL)在代码智能体中的应用通常使用可执行测试来提供二元奖励。基于这些奖励,组相对策略优化(GRPO)在每次 rollout 组内为通过测试的轨迹分配相同的优势,忽略了实现质量和任务需求遵循度的差异。这使得策略缺乏一种学习信号,无法偏好干净、有针对性的实现,而非那些包含不必要或超出范围更改的实现。我们提出了 GAGAR,一个用于代码智能体强化学习中质量感知的信用再分配框架。基于动态采样,保留同时包含通过和失败轨迹的组,GAGAR 将每组中的所有轨迹置于一个共享工作空间中,由一个经 SFT 训练的智能体评分器联合检查它们并对通过测试的候选者进行排名。基于该排名,我们对排名较低的轨迹进行降权,并按比例重新缩放所有通过测试轨迹的优势,以恢复其原始总和。这种保持总和的再分配保留了基于质量的降权所建立的相对权重,同时将信用转移至更高质量的实现。我们使用 MiMo-V2.6-Flash(总参数 310B)和 MiMo-V2.6-Pro(总参数 1.02T)的预 RL SFT 检查点,在工业规模上评估了 GAGAR。受控的仅代码 Flash 实验显示了改进的代码智能体性能、减少的轨迹长度增长和更稳定的训练。我们进一步将 GAGAR 应用于大规模混合任务 RL,同时使用 Flash 和 Pro。我们的结果支持将基于测试的验证与分组智能体评分相结合,以提高代码智能体强化学习的质量和稳定性。
英文摘要:
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate GAGAR at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply GAGAR in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.