发表机构
Rochester Institute of Technology; Adobe Research(罗切斯特理工学院; 奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现GRPO存在虚假优势会误导策略走向猜测行为,提出SIGNBALANCE算法,在有界答案数学任务和搜索智能体上性能优于GRPO,开放答案数学任务表现相当。
AI 中文摘要
组相对策略优化(GRPO)是针对可验证奖励的强化学习被广泛研究的算法,其优势估计器会根据组内奖励统计为每个rollout分配一个量级。在常见情况下,该量级会奖励那些通过推理得出正确答案的rollout。然而,存在一种被忽视的情况:某个rollout可能是通过猜测得到正确答案的,而公式仍会为其分配高量级,我们将此识别为虚假优势。这种情况出现在三类场景中:候选集较小的有界答案任务;包含有界子情况的开放答案集;以及预算为同一答案开辟多条路径的搜索智能体。在所有这三类场景中,虚假优势会误导策略走向类似猜测的行为。我们提出SIGNBALANCE,其量级是无组合的:它保留验证器符号、使用全局尺度,并通过逐类重缩放的停止梯度恢复零均值平衡。在不同规模的数学和搜索智能体基准测试中,SIGNBALANCE在开放答案数学任务上与GRPO表现相当,且在有界答案数学任务和搜索智能体上优于GRPO。代码将被发布。
英文摘要
Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable rewards, where its advantage estimator assigns each rollout a magnitude from within-group reward statistics. In the common case, this magnitude rewards rollouts that reach the correct answer through reasoning. Yet, an overlooked case shares the same surface: a rollout may land on it by guessing, and the formula still assigns a high magnitude, which we identify as the spurious advantage. This arises in three cases: bounded-answer tasks with a small candidate set; open-answer sets hosting bounded sub-cases; and search agents whose budget opens many paths to the same answer. In all three, this misleads the policy toward guess-like behaviors. We propose SIGNBALANCE, whose magnitude is composition-free: it keeps the verifier sign, uses a global scale, and restores zero-mean balance via a stop-gradient per-class rescaling. Across math and search agent benchmarks at different scales, SIGNBALANCE matches GRPO on open-answer math and improves on bounded-answer math and search agents. Code will be released.