发表机构
Graduate School of Data Science, Seoul National University(首尔国立大学数据科学研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对GRPO过度关注高概率响应导致推理覆盖度下降的问题,提出ReCo重加权方法,在多个数学推理基准上提升了大k值下的Pass@k表现。
AI 中文摘要
Group Relative Policy Optimization(GRPO)已成为后训练语言模型的标准强化学习方法。近期研究表明,GRPO会降低基础模型的推理能力,当k较大时在Pass@k指标上表现不如基础模型,这说明其推理路径覆盖度有所下降。我们发现这种性能下降与GRPO过度关注基础模型已以高概率生成的响应有关,我们将这种集中现象归因于GRPO更新中的两种机制:在响应层面,高概率响应因重复出现而主导群体梯度;在token层面,GRPO的重要性比率会缩放梯度,进一步强化当前策略下更可能出现的token。我们提出ReCo,一种可同时解决上述两种效应的重加权方法:响应贡献按其在rollout群体内的预期出现次数进行归一化,且将token层面的重要性比率替换为基于方差的比率,该比率会对非饱和决策点(即仍存在合理备选token选择的决策点)赋予更大的更新规模。在Qwen2.5-Math-1.5B/7B和Llama-3.1-8B-Instruct这两个模型、五个数学推理基准测试上,ReCo在k值较大时提升了Pass@k指标,在k值较小时与GRPO表现相当。
英文摘要
Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths. We find that this reduction is associated with GRPO concentrating on responses that the base model already generates with high probability. We trace this concentration to two mechanisms in the GRPO update. At the response level, high-probability responses dominate the group gradient through repeated occurrence. At the token level, GRPO's importance ratio scales gradients, further reinforcing tokens that become more likely under the current policy. We propose ReCo, a reweighting method that addresses both effects. Response contributions are normalized by their expected occurrence within the rollout group, and the token-level importance ratio is replaced with a variance-based ratio that gives larger update scale to non-saturated decision points where alternative token choices remain plausible. Across Qwen2.5-Math-1.5B/7B and Llama-3.1-8B-Instruct on five mathematical reasoning benchmarks, ReCo improves Pass@k for large values of k and is comparable to GRPO for small values of k.