发表机构
Meta AI; Columbia University(Meta AI; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用特权信息解决强化学习难题,提出脱离上下文的广义信赖域策略优化算法(OC - GRPO),通过引导展开和重要性校正目标避免训练不匹配,在标准数学推理基准上比普通GRPO有显著提升且成本低。
AI 中文摘要
具有可验证奖励的强化学习(RLVR)提升了大语言模型中的推理能力。然而,典型的RLVR方法在难题上会失效,当模型无法生成正确解决方案时,它会收到零学习信号。在训练期间提供特权指导,如解决方案前缀,可通过引导模型走向有非零奖励的正确解决方案来克服这一学习困境。我们将这些展开称为“脱离上下文”,它们由包含特权指导的训练提示生成,而目标目标由没有该指导的原始提示定义。我们引入了脱离上下文的广义信赖域策略优化算法(OC - GRPO),它是GRPO的一个最小修改变体,使用引导展开但应用重要性校正目标来引导更新回到原始无引导目标,避免了使未校正引导训练不稳定的不匹配。通过实验,我们的算法在标准数学推理基准上平均比普通GRPO实现了3.9%的绝对提升(13.8%的相对增益),且额外成本可忽略不计。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning signal. Providing privileged guidance during training, such as solution prefixes, can help overcome this learning cliff by steering the model towards {correct solutions with non-zero reward}. {We call these rollouts \textit{off-context}: they are generated from a training prompt that contains privileged guidance, while the target objective is defined by the original prompt without that guidance.} {We introduce} Off-Context GRPO (OC-GRPO), a minimally modified variant of GRPO that uses guided rollouts but applies an importance-corrected objective to steer the update back toward the original unguided objective, avoiding the mismatch that destabilizes uncorrected guided training. Empirically, our algorithm achieves a 3.8\% absolute improvement (13.7\% relative gain) over vanilla GRPO on average across standard mathematical reasoning benchmarks with negligible additional cost.
Comments24 Pages