AI 中文总结
该研究针对在线策略自蒸馏的暴露偏差问题,提出CRPO方法,通过对比学习与分组对比保留优化信号,在13个推理基准上提升了训练稳定性与泛化性。
AI 中文摘要
近期预训练后大语言模型(LLMs)的进展越来越依赖可验证奖励强化学习(RLVR)或在线策略自蒸馏(OPSD)。OPSD提供密集的logit级监督,但因自教师的特权信息固有存在暴露偏差,在多轮智能体设置中会导致推理路径收敛、丢失明确优化方向。为应对这些挑战,我们提出对比强化策略优化(CRPO),从对比学习视角重构智能体OPSD,利用预测熵区分反映探索的正样本位置与暴露偏差的负样本位置,通过分组对比保留可靠细粒度优化信号。在13个具挑战性的推理与深度搜索基准上的广泛评估表明,CRPO始终优于现有强化学习和自蒸馏基线,显著提升长程交互中的训练稳定性与泛化性。
英文摘要
Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level supervision, it inherently suffers from exposure bias due to the privileged information of the self-teacher. In multi-turn agentic settings, this leads to reasoning route convergence and the loss of clear optimization directions. To tackle these challenges, we introduce Contrastive Reinforced Policy Optimization (CRPO), which reformulates agentic OPSD from a contrastive learning perspective. By leveraging predictive entropy to distinguish between positive positions (reflective exploration) and negative positions (exposure bias), CRPO conducts group-wise contrast to preserve reliable, fine-grained optimization signals. Extensive evaluations across 13 challenging reasoning and deep-search benchmarks demonstrate that CRPO consistently outperforms existing reinforcement learning and self-distillation baselines, significantly enhancing training stability and generalization in long-horizon interactions.