arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RRPO:基于分层条件展开的参考相对策略优化

RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts

Yuxin Xiong, Xunyi Jiang, Rohan Surana, Xintong Li, Sheldon Yu, Nikki Lijing Kuang, Ryan A. Rossi, Jingbo Shang, Tong Yu, Julian McAuley, Junda Wu

arXiv 2607.18470首次发表:更新:

发表机构

University of California San Diego; Adobe Research(加利福尼亚大学圣地亚哥分校; Adobe研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出RRPO,通过参考相对对比比较推广GRPO。用分层条件展开构建锚集,训练投影头定义对比优势,在策略优化中评估。不依赖任务真实验证器,在多种设置下有竞争力,优于弱监督基线,监督微调后有额外收益。

AI 中文摘要

组相对策略优化(GRPO)在可验证反馈的强化学习中已显示出强大的有效性,其中可以使用任务提供的正确性信号在组内比较采样的展开。然而,将组相对优化扩展到可验证设置之外具有挑战性,因为许多任务的成功无法通过单一正确性标准来衡量。我们提出了参考相对策略优化(RRPO),它通过用参考相对对比比较取代基于直接正确性的优势构建来推广GRPO。RRPO首先使用分层条件展开来构建正锚集和负锚集,然后使用集对比目标训练一个度量投影头,以将候选展开与这些锚进行比较。由此产生的对齐分数直接定义对比优势:在策略优化期间,投影头被冻结,分数在标准组相对目标中的每个展开组内居中。我们在整个策略优化过程中使用基于锚的对比优势来评估RRPO,而不依赖于任务真实验证器。在可验证推理、开放式生成和后SFT设置中,RRPO与基于验证器的优化相比仍具有竞争力,优于弱监督基线,并在监督微调后提供额外收益。

英文摘要

Group Relative Policy Optimization (GRPO) has shown strong effectiveness in reinforcement learning from verifiable feedback, where sampled rollouts can be compared within a group using task-provided correctness signals. However, extending group-relative optimization beyond verifiable settings is challenging because success in many tasks is not captured by a single correctness criterion. We propose \textbf{Reference-Relative Policy Optimization (RRPO)}, which generalizes GRPO by replacing direct correctness-based advantage construction with reference-relative contrastive comparisons. RRPO first uses \emph{stratified conditional rollouts} to construct positive and negative anchor sets, and then trains a metric projection head with a set-contrastive objective to compare candidate rollouts against these anchors. The resulting alignment scores directly define contrastive advantages: during policy optimization, the projection head is frozen, and the scores are centered within each rollout group in a standard group-relative objective. We evaluate RRPO using anchor-based contrastive advantages throughout policy optimization, without relying on task ground-truth verifiers. Across verifiable reasoning, open-ended generation, and post-SFT settings, RRPO remains competitive with verifier-based optimization, improves over weakly supervised baselines, and provides additional gains after supervised fine-tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑