AI 中文总结
针对GRPO中组梯度冲突导致策略更新效果差的问题,提出GUPO方法,通过贝叶斯框架建模组梯度并结合狄利克雷公式校准梯度贡献,经多基准实验验证其有效性。
AI 中文摘要
组相对策略优化(GRPO)已成为后训练大语言模型(LLMs)推理的广泛应用方法。在GRPO中,同一小批量内不同查询产生的组梯度被直接平均以形成策略更新,但这些组梯度可能指向冲突方向。我们的实证分析表明,组梯度冲突往往与效果较差的策略更新相关,这促使我们需要在这种冲突下找到可靠的聚合更新方向。标准GRPO聚合将实际组梯度视为确定性贡献,未考虑聚合过程中它们的可靠性差异。为解决此问题,我们提出梯度不确定性感知策略优化(GUPO),该方法在贝叶斯框架下将每个组梯度建模为随机变量并估计其概率分布,随后基于狄利克雷(Dirichlet)公式推导梯度不确定性,并用其在校准聚合过程中每个组梯度的贡献。在多个基准上的大量实验证明了GUPO的有效性。
英文摘要
Group Relative Policy Optimization (GRPO) has become a widely used approach for post-training Large Language Models (LLMs) for reasoning. In GRPO, the group gradients induced by different queries within the same mini-batch are directly averaged to form the policy update. However, these group gradients can point in conflicting directions. Our empirical analysis suggests that group-gradient conflicts tend to be associated with less effective policy updates, motivating the need for a reliable aggregated update direction under such conflicts. Standard GRPO aggregation treats the realized group gradients as deterministic contributions and does not account for differences in their reliability during aggregation. To address this issue, we propose Gradient Uncertainty-Aware Policy Optimization (GUPO), which models each group gradient as a random variable under a Bayesian formulation and estimates its probability distribution. GUPO then derives gradient uncertainty using a Dirichlet-based formulation and uses it to calibrate the contribution of each group gradient during aggregation. Extensive experiments on multiple benchmarks demonstrate the effectiveness of GUPO.