发表机构
Renmin University of China; Tsinghua University; University of Electronic Science and Technology of China; Shanghai Jiao Tong University(中国人民大学; 清华大学; 电子科技大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TIPS框架,通过仅结果强化学习诱导过程监督,生成式PRM以低成本训练,在数学和智能体基准上超越现有PRMs和强提示判断器。
AI 中文摘要
过程奖励模型(PRMs)已成为大型语言模型(LLMs)的关键组成部分,因为其步骤级反馈支持训练后推理和测试时推理。然而,训练强大的PRMs仍然代价高昂:人工步骤标注难以扩展,而蒙特卡洛估计计算成本高且可能偏离步骤的内在正确性。为了以低成本获得有效的PRMs,我们引入了TIPS(思考诱导的过程监督),一个仅结果的强化学习(RL)框架,用于训练生成式PRMs。在TIPS中,模型生成思维链(CoT),随后生成步骤级标签和结果标签。奖励仅取决于预测结果是否与真实结果匹配,由此产生的组相对优势用于优化整个生成的响应。直观上,当检查中间步骤有助于确定结果时,更准确的检查可以导致更好的结果判断和更高的奖励。因此,仅结果的强化学习可以在没有显式过程监督的情况下强化步骤级验证。我们在数学和智能体基准以及四个骨干家族上验证了TIPS的有效性。值得注意的是,TIPS-Qwen3-4B-Thinking-2507在ProcessBench上仅使用3.2K个结果标记的轨迹就达到了85.2的F1分数,超过了所有评估的训练PRMs和强大的仅提示判断器,如GPT-5.4-Instruct和Claude-4.7-Opus,但仍落后于o1-mini。代码和数据可在该https URL获取。
英文摘要
Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning. However, training strong PRMs remains costly: human step annotation is difficult to scale, while Monte Carlo estimation is computationally expensive and can drift from the intrinsic correctness of steps. To get effective PRMs at low cost, we introduce TIPS (Thinking-Induced Process Supervision), an outcome-only reinforcement learning (RL) framework for training generative PRMs. In TIPS, the model generates a chain-of-thought (CoT) followed by step-level labels and an outcome label. The reward depends solely on whether the predicted outcome matches the ground truth, and the resulting group-relative advantage is used to optimize the entire generated response. Intuitively, when checking intermediate steps helps determine the outcome, more accurate checks can lead to better outcome judgments and higher rewards. Outcome-only RL can therefore reinforce step-level verification without explicit process supervision. We validate the effectiveness of TIPS across math and agent benchmarks and four backbone families. Notably, TIPS-Qwen3-4B-Thinking-2507 reaches 85.2 F1 on ProcessBench with only 3.2K outcome-labeled trajectories, surpassing all evaluated trained PRMs and strong prompt-only judges such as GPT-5.4-Instruct and Claude-4.7-Opus, while still trailing o1-mini. Code and data are available at https://github.com/RUCBM/TIPS.