发表机构
University of California San Diego; Adobe Research(加利福尼亚大学圣地亚哥分校; Adobe研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出RRPO,通过参考相对对比比较推广GRPO。用分层条件展开构建锚集,训练投影头定义对比优势,在策略优化中评估。不依赖任务真实验证器,在多种设置下有竞争力,优于弱监督基线,监督微调后有额外收益。
AI 中文摘要
组相对策略优化(GRPO)在可验证反馈的强化学习中已显示出强大的有效性,其中可以使用任务提供的正确性信号在组内比较采样的展开。然而,将组相对优化扩展到可验证设置之外具有挑战性,因为许多任务的成功无法通过单一正确性标准来衡量。我们提出了参考相对策略优化(RRPO),它通过用参考相对对比比较取代基于直接正确性的优势构建来推广GRPO。RRPO首先使用分层条件展开来构建正锚集和负锚集,然后使用集对比目标训练一个度量投影头,以将候选展开与这些锚进行比较。由此产生的对齐分数直接定义对比优势:在策略优化期间,投影头被冻结,分数在标准组相对目标中的每个展开组内居中。我们在整个策略优化过程中使用基于锚的对比优势来评估RRPO,而不依赖于任务真实验证器。在可验证推理、开放式生成和后SFT设置中,RRPO与基于验证器的优化相比仍具有竞争力,优于弱监督基线,并在监督微调后提供额外收益。
英文摘要
Group Relative Policy Optimization (GRPO) has shown strong effectiveness in reinforcement learning from verifiable feedback, where sampled rollouts can be compared within a group using task-provided correctness signals. However, extending group-relative optimization beyond verifiable settings is challenging because success in many tasks is not captured by a single correctness criterion. We propose \textbf{Reference-Relative Policy Optimization (RRPO)}, which generalizes GRPO by replacing direct correctness-based advantage construction with reference-relative contrastive comparisons. RRPO first uses \emph{stratified conditional rollouts} to construct positive and negative anchor sets, and then trains a metric projection head with a set-contrastive objective to compare candidate rollouts against these anchors. The resulting alignment scores directly define contrastive advantages: during policy optimization, the projection head is frozen, and the scores are centered within each rollout group in a standard group-relative objective. We evaluate RRPO using anchor-based contrastive advantages throughout policy optimization, without relying on task ground-truth verifiers. Across verifiable reasoning, open-ended generation, and post-SFT settings, RRPO remains competitive with verifier-based optimization, improves over weakly supervised baselines, and provides additional gains after supervised fine-tuning.