解锁评论家:面向LLM后训练的无奖励策略优化
Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training
浏览论文内容
中文总结 AI 辅助
本文提出无奖励策略优化(RFPO),利用预训练冻结评论家提供密集逐前缀信号,无需标签即可匹配监督式PPO,并显著降低计算与内存开销,适用于长程推理任务。
中文摘要 AI 辅助
近期针对大型语言模型的强化学习(RL)后训练方法日益倾向于移除评论家(critic)以减少训练不稳定性和内存开销。即使在训练过程中使用了评论家,训练结束后也会将其丢弃,尽管它已经学会了预测结果。我们重新审视这一趋势,并表明预训练评论家预测未来结果的能力可使其成为高效长程推理的宝贵资产。首先,我们发现基于评论家的RL在长思维链推理中的不稳定性在很大程度上是一个优化伪影:保持策略更新幅度小且方差低即可恢复稳定收敛。其次,一个训练良好的评论家能够从后续轨迹状态和未完成前缀中估计最终成功的后验概率。其预测提供了基于结果、密集的逐前缀学习信号,在策略优化过程中既不需要完整的轨迹展开,也不需要逐步标注,更不需要外部奖励标签。基于这一见解,我们引入了无奖励策略优化(RFPO),该方法将单一校准的冻结评论家重新用作轨迹级奖励、用于广义优势估计的值基线,以及用于未完成前缀的成功预测器。我们进一步表明,对去偏分数进行二值化可阻止策略利用评论家的长度偏差。二值化后的RFPO在训练循环中无需任何标签即可匹配监督式PPO的性能,同时削减了计算和内存开销。这使得RFPO非常适合长程推理任务,此类任务中结果出现较晚且生成过程主导成本:由于轨迹在完成前即可获得奖励,训练不再需要为等待每条轨迹完成而付出代价。我们的研究结果挑战了当前无评论家范式,并确立了基于评论家的无奖励优化作为LLM后训练的一种可扩展且计算高效的路径。
英文摘要
Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning. First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small and low in variance restores stable convergence. Second, a well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. Its predictions provide outcome-derived, dense, per-prefix learning signals that, during policy optimization, require neither completed rollouts, step-level annotations, nor external reward labels. Building on this insight, we introduce Reward-Free Policy Optimization (RFPO), which repurposes a single calibrated, frozen critic as a rollout-level reward, a value baseline for generalized advantage estimation, and a success forecaster for unfinished prefixes. We further show that binarizing the debiased score stops the policy from exploiting the critic's length bias. Binarized, RFPO matches supervised PPO without a single label in the training loop, while cutting compute and memory overhead. This makes RFPO well suited to long-horizon reasoning tasks, where outcomes arrive late and generation dominates cost: because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete. Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.
发表机构
- University of Luxembourg(卢森堡大学)
- Seafill Open-Source Community(Seafill 开源社区)
- Université Paris-Saclay(巴黎-萨克雷大学)
机构由 AI 辅助整理,请以论文原文为准。