发表机构
Zhejiang University; ByteDance(浙江大学; 字节跳动)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于评分标准的强化学习中探索有限问题,提出标准蒸馏策略优化(CriPO),通过策略内自蒸馏同时解决未探索标准和被抑制标准问题,在医学和科学基准测试中表现出色,减少优化步骤并提升性能。
AI 中文摘要
基于评分标准的强化学习在改进大语言模型的开放式任务中展现出潜力。其公认的局限是探索有限,未探索标准(UC)未获优化信号。近期方法通过在展开过程中纳入评分标准信息解决此问题,但引入了训练与推理不匹配。此外,这些方法忽视了被抑制标准(SC)这一失败模式。我们的分析表明SC很普遍。为同时解决UC和SC且不引入训练推理不匹配,我们提出标准蒸馏策略优化(CriPO),通过策略内自蒸馏增强基于评分标准的强化学习。在医学和科学基准测试中,CriPO持续优于基于评分标准的强化学习,以约少两倍的优化步骤实现更强的最终性能。
英文摘要
Rubric-based Reinforcement Learning (RL) has recently shown promise in improving Large Language Models (LLMs) on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored Criteria) receive no optimization signal. Recent methods address this by incorporating rubric information as external guidance during rollout generation, yet they introduce a train-inference mismatch: the policy is optimized on rollouts produced under external guidance while this guidance is absent at inference time, causing error accumulation through autoregressive decoding. Moreover, these exploration-focused approaches overlook a fundamentally different failure mode that we term Suppressed Criteria---criteria that are satisfied by some rollouts yet whose learning signals might be lost during optimization because scalar reward aggregation assigns them non-positive aggregate advantages. Our analysis reveals that suppressed criteria constitute a persistent and non-negligible failure mode---over 25% of samples exhibit this issue throughout training. To simultaneously address both unexplored and suppressed criteria without introducing training-inference mismatch, we propose Criterion-Distilled Policy Optimization (CriPO), which enhances rubric-based RL via on-policy self-distillation. For unexplored criteria, CriPO constructs a behavior-injection teacher and computes a filtered forward-KL loss to inject missing behaviors into the policy. For suppressed criteria, CriPO uses a counterfactual teacher to locate criterion-relevant tokens in negative-advantage rollouts, and corrects their advantages in GRPO to preserve useful patterns. Experiments on medicine and science benchmarks demonstrate that CriPO outperforms existing rubric-based RL methods, e.g., achieving an average gain of 3.3 points over GRPO on Qwen3-4B.