ReDiPPO:用于数学推理的参考引导值校准与差异感知标记重加权
ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning
浏览论文内容
中文总结 AI 辅助
针对数学推理中PPO的标记级信用分配难题,ReDiPPO引入参考引导评判器并通过值估计差异重加权标记优势,提升值估计精度且优于PPO、DAPO等基线。
中文摘要 AI 辅助
强化学习已成为提升大语言模型数学推理能力的有效范式。在现有策略优化方法中,近端策略优化(Proximal Policy Optimization, PPO)颇具吸引力,因为其学习到的评判器原则上可提供标记级信用分配。但在具有长推理视界和稀疏结果奖励的数学推理任务中,可靠的标记级信用分配仍具挑战性。标准评判器常无法准确评估中间推理状态,导致优势估计噪声及次优策略更新。本文提出ReDiPPO,一种用于数学推理的参考引导且差异感知的PPO框架。ReDiPPO引入参考引导评判器,利用参考答案作为训练时的特权信号提供更准确的值估计;同时保留标准评判器,量化标准值估计与参考引导值估计间的标记级参考-标准差异,该差异作为困难推理状态的指标,用于在PPO优化期间对相应标记级优势进行重加权。在多种数学推理基准上的大量实验表明,ReDiPPO提升了值估计准确性,且在最终推理性能上始终优于PPO、DAPO和GSPO等强策略优化基线。我们的代码可在该httpsURL获取。
英文摘要
Reinforcement learning has emerged as an effective paradigm for enhancing the mathematical reasoning capabilities of large language models. Among existing policy optimization methods, Proximal Policy Optimization (PPO) remains particularly appealing because its learned critic can, in principle, provide token-level credit assignment. However, in mathematical reasoning tasks characterized by long reasoning horizons and sparse outcome rewards, reliable token-level credit assignment remains challenging. The standard critic often fails to accurately evaluate intermediate reasoning states, resulting in noisy advantage estimates and suboptimal policy updates. In this paper, we propose ReDiPPO, a Reference-guided and Discrepancy-aware PPO framework for mathematical reasoning. ReDiPPO introduces a reference-guided critic that uses reference answers as training-time privileged signals to provide more accurate value estimation. Meanwhile, it retains a standard critic and quantifies the token-level reference-standard discrepancy between the standard value estimate and the reference-guided value estimate. This discrepancy serves as an indicator of difficult reasoning states and is used to reweight the corresponding token-level advantages during PPO optimization. Extensive experiments on diverse mathematical reasoning benchmarks demonstrate that ReDiPPO improves value-estimation accuracy and consistently outperforms strong policy optimization baselines, including PPO, DAPO, and GSPO, in final reasoning performance. Our code is available on https://github.com/cii030/ReDiPPO.