arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32577cs.CLcs.LG

分组智能体评分与优势再分配用于代码智能体强化学习

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

AI总结:

提出GAGAR框架,通过智能体评分器对组内通过测试的轨迹进行排名并重新分配优势,提升代码智能体强化学习的质量与稳定性。

AI中文摘要:

强化学习(RL)在代码智能体中的应用通常使用可执行测试来提供二元奖励。基于这些奖励,组相对策略优化(GRPO)在每次 rollout 组内为通过测试的轨迹分配相同的优势,忽略了实现质量和任务需求遵循度的差异。这使得策略缺乏一种学习信号,无法偏好干净、有针对性的实现,而非那些包含不必要或超出范围更改的实现。我们提出了 GAGAR,一个用于代码智能体强化学习中质量感知的信用再分配框架。基于动态采样,保留同时包含通过和失败轨迹的组,GAGAR 将每组中的所有轨迹置于一个共享工作空间中,由一个经 SFT 训练的智能体评分器联合检查它们并对通过测试的候选者进行排名。基于该排名,我们对排名较低的轨迹进行降权,并按比例重新缩放所有通过测试轨迹的优势,以恢复其原始总和。这种保持总和的再分配保留了基于质量的降权所建立的相对权重,同时将信用转移至更高质量的实现。我们使用 MiMo-V2.6-Flash(总参数 310B)和 MiMo-V2.6-Pro(总参数 1.02T)的预 RL SFT 检查点,在工业规模上评估了 GAGAR。受控的仅代码 Flash 实验显示了改进的代码智能体性能、减少的轨迹长度增长和更稳定的训练。我们进一步将 GAGAR 应用于大规模混合任务 RL,同时使用 Flash 和 Pro。我们的结果支持将基于测试的验证与分组智能体评分相结合,以提高代码智能体强化学习的质量和稳定性。

英文摘要:

Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate GAGAR at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply GAGAR in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.

↑