请勿偷看答案:用于无标签RLVR的结果掩码组相对策略优化
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
浏览论文内容
中文总结 AI 辅助
该研究针对无标签RLVR中模型易强化答案标记的问题,提出OM-GRPO框架,解耦奖励估计与策略优化,结合对比增强奖励,在多推理基准及测试时训练场景中性能优于现有方法。
中文摘要 AI 辅助
可验证奖励的强化学习(RLVR)可提升大型语言模型(LLM)的推理能力,但通常依赖真实(GT)答案,限制了可扩展性。基于投票的无标签RLVR用模型样本的答案级共识替代黄金监督。然而,当相同的答案级信号同时用于估计奖励和驱动标记级策略优化时,会出现崩溃问题,促使模型直接强化答案标记而非改进推理。我们提出OM-GRPO,一种无标签RLVR框架,将奖励估计与策略优化解耦。OM-GRPO掩码答案跨度的梯度,同时通过软共识信号保留答案级奖励,将优化压力从答案标记转移。我们进一步引入对比增强奖励,通过对现有轨迹进行低成本成对比较来优化奖励估计,无需额外的rollout。在多种推理基准和三种LLM骨干上,OM-GRPO始终优于现有无标签RLVR方法,且在稳定优化下可达到有监督GT奖励训练的效果。这种稳定性在测试时训练(Test-Time Training)场景中尤为有益,此时OM-GRPO超过多数投票方法4.24个百分点。
英文摘要
Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-level consensus from model samples. However, collapse arises when the same answer-level signal is used both to estimate rewards and to drive token-level policy optimization, encouraging the model to directly reinforce answer tokens rather than improve reasoning. We propose OM-GRPO, a label-free RLVR framework that decouples reward estimation from policy optimization. OM-GRPO masks gradients on the answer span while retaining answer-level rewards through a soft consensus signal, shifting optimization pressure away from answer tokens. We further introduce Contrast-Augmented Reward, which refines reward estimation via low-cost pairwise comparisons over existing trajectories without additional rollouts. Across diverse reasoning benchmarks and three LLM backbones, OM-GRPO consistently outperforms existing label-free RLVR methods and matches supervised GT-reward training with stable optimization. This stability is particularly beneficial in the Test-Time Training setting, where OM-GRPO surpasses majority voting by 4.24 points.