发表机构
Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多奖励GRPO中大规模相关奖励主导归一化的问题,提出CorrGRPO,将成对协方差归一化为皮尔逊相关系数,平衡不同尺度奖励影响,在代码生成、工具调用和智能体安全任务上取得改进。
AI 中文摘要
组相对策略优化(GRPO)被广泛用于训练推理语言模型,它通过对同一提示的多个回滚(rollouts)中的奖励进行中心化和归一化来计算优势。对于多个奖励,GRPO将各奖励分量求和,并用组内标准差对总奖励进行归一化。相应的方差等于所有成对奖励协方差之和。对于固定的中心化奖励,较大的总协方差会产生较小的优势,反之亦然,从而使更新幅度能够适应奖励之间的依赖性。然而,具有较大尺度的相关奖励可能主导这种归一化,并抑制来自较小尺度奖励的信号。我们提出了相关性归一化GRPO(CorrGRPO),它将成对协方差归一化为皮尔逊相关系数。CorrGRPO保持中心化总奖励不变,同时平衡不同尺度奖励对基于相关性的归一化的影响。这使得优势幅度能够适应奖励相关性,而不会被大规模奖励分量主导归一化。我们在代码生成、工具调用和智能体安全任务上,使用从0.5B到8B参数规模的模型,将CorrGRPO与GRPO及其他变体进行了比较。这些任务都涉及多个可以共同改进或呈现权衡的奖励。结果显示,在代码生成、工具调用和智能体安全这三个领域均有所改进。我们的代码可在以下网址获取:https://this URL。
英文摘要
Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards. We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization. This allows advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components. We compare CorrGRPO with GRPO and other variants on code generation, tool calling, and agent security, using models ranging from 0.5B to 8B parameters. These tasks all involve multiple rewards that can improve together or present tradeoffs. Results show improvements across three domains, including code generation, tool calling, and agent security. Our code is available at https://github.com/HKUST-KnowComp/CorrGRPO.