Verifiable Process Rewards for Agentic Reasoning
可验证过程奖励用于智能体推理
机构 * Tsinghua University(清华大学)
专题命中 推理与问题求解 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 提出可验证过程奖励(VPR)框架,通过符号或算法预言机将密集的中间步骤验证转化为强化学习的逐轮监督信号,解决长程智能体推理中的信用分配问题,并在搜索、约束和后验验证三种场景中验证其有效性。
Comments Corrected minor typos and LLM-assisted data extraction errors. The main conclusions are unchanged