发表机构
The Chinese University of Hong Kong; South China University of Technology; Nanyang Technological University(香港中文大学; 华南理工大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出对比策略优化(CPO),利用参考引导与普通生成分布的令牌级对比分歧实现正确性感知优势塑造,解决零优势问题,在基准实验中显著优于基于熵的方法,平衡探索与利用达最佳性能。
AI 中文摘要
具有可验证奖励的强化学习(RLVR)通常使用熵进行优势塑造。然而,熵无法区分有用的不确定性和有害的混淆,限制了其作为正确性信号的有效性。我们提出了对比策略优化(CPO),它利用参考引导和普通生成分布之间的令牌级对比分歧进行正确性感知优势塑造。理论和实证结果表明,这种分歧可靠地指示令牌级正确性。我们还表明,策略蒸馏是CPO的一种特殊情况,其中后验分布由外部教师模型实例化。CPO还解决了零优势问题。在域内和域外基准上的实验表明,CPO在保持强泛化能力的同时,显著优于基于熵的RLVR方法。进一步分析表明,正确和错误响应分别自然地支持探索和利用,平衡两者可导致最佳性能。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We propose Contrastive Policy Optimization (CPO), which uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping. Both theoretical and empirical results show that this disagreement reliably indicates token-level correctness. We further show that On-policy Distillation is a special case of CPO, where the posterior distribution is instantiated by an external teacher model. CPO also resolves the zero-advantage problem. Experiments on in-domain and out-of-domain benchmarks demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Further analysis shows that correct and incorrect responses naturally support exploration and exploitation respectively, and balancing both leads to the best performance.