发表机构
KAIST; AITRICS(韩国科学技术院; AITRICS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对RLVR中全失败组缺乏策略梯度信号的问题,提出离策略感知的GRAFT框架,通过跨模型交换同伴轨迹并控制不匹配,在相同预算下提升异构模型性能,平均增益2.1点。
AI 中文摘要
基于可验证奖励的强化学习(RLVR)方法,如GRPO,依赖于成功的自生成轨迹,但有限的rollout预算可能产生全失败组,导致没有基于奖励的策略梯度信号。虽然增加rollout次数能以更高成本提高成功概率,但一个模型rollout中缺失的成功轨迹可能已被另一个模型发现。实际上,我们观察到异构模型往往在互补的提示上成功,这为无需指定更强教师的情况下进行相互学习创造了机会。为利用这种互补性,我们提出GRAFT(用同伴轨迹门控替换回答失败组),一种离策略感知框架,用信息丰富的同伴组替换全失败组。GRAFT转移成功和不成功的同伴响应及其同伴计算的优势,同时通过序列级兼容性加权和令牌级重要性比裁剪来控制跨模型不匹配。在三个异构模型对和五个数学推理基准上,GRAFT在相同每模型rollout预算下,相较于GRPO持续改进两个模型,平均提升2.1个百分点,模型级平均性能最高提升4.5个百分点。存储的同伴轨迹保留了大部分增益,在没有同时协同训练的情况下,平均比GRPO提升1.8个百分点。
英文摘要
Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups. GRAFT transfers both successful and unsuccessful peer responses with peer-computed advantages, while controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping. Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.
Comments29 pages, 11 figures, 9 tables