SoftmaxGRPO:使用Softmax优势组估计学习推理
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation
浏览论文内容
中文总结 AI 辅助
该研究针对GRPO在二元奖励下对简单提示加权发散的问题,提出SoftmaxGRPO替换其归一化方式,经实证该方法可重新分配梯度预算,在DeepMath和Poetry任务上均优于GRPO并取得显著性能提升。
中文摘要 AI 辅助
基于组的强化学习目标(如GRPO)在提示难度间分配学习信号的效果较差:在二元奖励下,组归一化会对简单提示产生发散的加权。我们引入Softmax优势组估计(SoftmaxGRPO),这是一种即插即用的替代方案,它将z分数归一化的组优势替换为温度缩放的softmax优势,无论提示难度如何都能保持权重有界。对于二元奖励,我们推导了精确的有限组总体目标,并确定MaxRL是其低温极限。对于有界标量奖励,我们表明大组更新恰好优化对数矩生成函数目标,而通用有限组标量目标在不对奖励分布做额外假设的情况下不存在。实证结果显示,SoftmaxGRPO将测量的梯度预算从接近解决的提示重新分配,在相同奖励下始终优于GRPO;它在具有可验证奖励的DeepMath上达到51.8%,仅使用轻量级文本相似性奖励就将15亿参数的指令调优模型在Poetry上的性能从35.0%提升至68.0%。
英文摘要
Group-based reinforcement learning objectives such as GRPO can allocate learning signal poorly across prompt difficulty: under binary rewards, group normalization induces a divergent weighting on easy prompts. We introduce Softmax Advantage Group Estimation (SoftmaxGRPO), a drop-in alternative that replaces z-score-normalized group advantages with temperature-scaled softmax advantages, keeping weights bounded regardless of prompt difficulty. For binary rewards, we derive the exact finite-group population objective and identify MaxRL as its low-temperature limit. For bounded scalar rewards, we show that the large-group update exactly optimizes a log-moment-generating-function objective, while a universal finite-group scalar objective cannot exist without additional assumptions on the reward distribution. Empirically, SoftmaxGRPO reallocates measured gradient budget away from near-solved prompts and consistently improves over GRPO under identical rewards. It reaches 51.8% on DeepMath with verifiable rewards and improves a 1.5B instruction-tuned model from 35.0% to 68.0% on Poetry using only lightweight text-similarity rewards.
发表机构
- Rice University(莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。