可验证过程奖励用于智能体推理
Verifiable Process Rewards for Agentic Reasoning
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出可验证过程奖励(VPR)框架,通过符号或算法预言机将密集的中间步骤验证转化为强化学习的逐轮监督信号,解决长程智能体推理中的信用分配问题,并在搜索、约束和后验验证三种场景中验证其有效性。
AI中文摘要:
来自可验证奖励的强化学习(RLVR)提升了大型语言模型(LLMs)的推理能力,但现有方法大多依赖稀疏的结果级反馈。这种稀疏性在长程智能体推理中造成了信用分配难题:一个轨迹可能因包含许多正确的中间决策而失败,或因包含有缺陷的决策而成功。在这项工作中,我们研究了一类密集可验证的智能体推理问题,其中中间动作可以通过符号或算法预言机进行客观检查。我们提出了可验证过程奖励(VPR),一个将此类预言机转化为密集的逐轮监督信号用于强化学习的框架,并在三个代表性场景中实例化:基于搜索验证的动态演绎、基于约束验证的逻辑推理和基于后验验证的概率推理。我们进一步提供了理论分析,表明密集的验证器基础奖励可以通过提供更局部的学习信号来改善长程信用分配,其收益取决于验证器的可靠性。实验上,VPR在受控环境中优于结果级奖励和基于rollout的过程奖励基线,更重要的是,它能够迁移到通用和智能体推理基准测试中,表明可验证的过程监督可以培养适用于训练环境之外的通用推理技能。我们的结果表明,VPR是一种有前景的方法,用于在可靠的中间验证可用时增强LLM智能体,同时也强调了其对预言机质量的依赖性,以及将VPR扩展到结构化程度较低、开放环境中的开放性挑战。
英文摘要:
Reinforcement learning from verifiable rewards (RLVR) has improved the reasoning abilities of large language models (LLMs), but most existing approaches rely on sparse outcome-level feedback. This sparsity creates a credit assignment challenge in long-horizon agentic reasoning: a trajectory may fail despite containing many correct intermediate decisions, or succeed despite containing flawed ones. In this work, we study a class of densely-verifiable agentic reasoning problems, where intermediate actions can be objectively checked by symbolic or algorithmic oracles. We propose Verifiable Process Rewards (VPR), a framework that converts such oracles into dense turn-level supervision for reinforcement learning, and instantiate it in three representative settings: search-based verification for dynamic deduction, constraint-based verification for logical reasoning, and posterior-based verification for probabilistic inference. We further provide a theoretical analysis showing that dense verifier-grounded rewards can improve long-horizon credit assignment by providing more localized learning signals, with the benefit depending on the reliability of the verifier. Empirically, VPR outperforms outcome-level reward and rollout-based process reward baselines across controlled environments, and more importantly, transfers to both general and agentic reasoning benchmarks, suggesting that verifiable process supervision can foster general reasoning skills applicable beyond the training environments. Our results indicate that VPR is a promising approach for enhancing LLM agents whenever reliable intermediate verification is available, while also highlighting its dependence on oracle quality and the open challenge of extending VPR to less structured, open-ended environments.