发表机构
Southeast University; Key Laboratory of Computer Network and Information Integration (Southeast University); Ministry of Education, China; Zhongguancun Academy; Zhongguancun Institute of Artificial Intelligence(东南大学; 计算机网络和信息集成重点实验室(东南大学); 中国教育部; 中关村学院; 中关村人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出推理状态传播(RSP)方法,通过建模推理状态转移连接中间与最终状态,利用结果监督辅助过程学习,在推理搜索、响应选择和强化学习中显著优于现有PRM基线。
AI 中文摘要
过程奖励模型(PRMs)通过在测试时扩展和强化学习中提供评估中间推理状态的细粒度信号,展现出显著的有效性,但其训练严重依赖昂贵的过程标注。缓解这种依赖的一种自然方式是用可扩展的结果监督来补充有限的过程监督。然而,现有的PRM通常独立地对推理前缀进行建模,缺乏有效利用最终结果来指导中间推理状态学习的显式机制。我们提出了推理状态传播(RSP),该方法将每个推理前缀表示为一个二值有效状态,并对推理轨迹中连续状态之间的转移进行建模。具体而言,RSP预测一个有效状态变为无效的破坏概率,以及一个无效状态恢复为有效的修复概率。通过传播这些转移,RSP将中间状态与最终状态连接起来,使得过程标注可以监督中间状态,而结果标签则监督最终状态,并能向前面的步骤提供学习信号。在推理搜索、响应选择和强化学习中,RSP consistently优于代表性的PRM基线,在束搜索中相比Qwen2.5-Math-PRM平均提升5.6%,在强化学习中提升2.1%。
英文摘要
Process reward models (PRMs) have demonstrated notable effectiveness in test-time scaling and reinforcement learning by providing fine-grained signals for evaluating intermediate reasoning states, but their training relies heavily on costly process annotations. A natural way to alleviate this dependence is to complement limited process supervision with scalable outcome supervision. However, existing PRMs often model reasoning prefixes independently, providing no explicit mechanism for effectively using final outcome to guide the learning of intermediate reasoning states. We introduce Reasoning State Propagation (RSP), which represents each reasoning prefix with a binary validity state and models transitions between successive states across the reasoning trajectory. Specifically, RSP predicts a break probability that a valid state becomes invalid and a repair probability that an invalid state returns to valid. By propagating these transitions, RSP connects intermediate states to the final state, allowing process annotations to supervise intermediate states while outcome labels supervise the final state and can provide learning signals to preceding steps. Across reasoning search, response selection, and reinforcement learning, RSP consistently outperforms representative PRM baselines, with average improvements over Qwen2.5-Math-PRM of 5.6% in beam search and 2.1% in reinforcement learning.