arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期望推理步回报统一了基于策略的奖励与教师学习

Expected Reasoning-Step Return Unifies On-Policy Learning from Rewards and Teachers

Qiangqiang He, Jin Li

arXiv 2609.32674首次发表:更新:

AI 中文总结

本文提出期望推理步回报(ERSR)统一奖励与教师信号,并据此设计R²OPL算法,在成功轨迹强化学生推理、失败轨迹蒸馏教师信号,实验证明其优于现有基线。

AI 中文摘要

基于策略的推理模型可以从任务奖励或教师信号中学习,但这些来源在形式上不同,且可能偏向相互冲突的更新,导致不清楚在给定推理动作中应遵循哪个信号。我们引入了期望推理步回报(ERSR),将语义推理步骤视为宏动作,并使用蒙特卡洛学生策略展开来估计学生生成和教师提出的动作在共同回报空间中的期望最终任务奖励,以进行步骤级比较。ERSR分析揭示了一种结果依赖的不对称性:在成功轨迹上,学生动作比教师替换更有益,而在失败轨迹上,教师替换变得更有益。我们进一步表明,学生答案探测增益追踪学生步骤的ERSR效用,并区分有益和有害的推理步骤。基于这些发现,我们提出了回报引用的基于策略学习(R$^2$OPL),它在成功轨迹上强化学生推理,在失败轨迹上蒸馏教师信号,同时使用组成功率进行难度缩放,并使用学生探测增益进行步骤级调制。在推理基准和教师-学生配置上的实验表明,R$^2$OPL始终优于强基线。ERSR训练动态进一步表明,R$^2$OPL联合利用了奖励侧和教师侧信号的实质性效用,而现有的混合方法往往在一个分支中留下大量残余效用。

英文摘要

On-policy reasoning models can learn from task rewards or teacher signals, but these sources differ in form and can favor conflicting updates, leaving unclear which should guide a given reasoning action. We introduce \textbf{Expected Reasoning-Step Return (ERSR)}, which treats semantic reasoning steps as macro-actions and uses Monte Carlo student-policy rollouts to estimate the expected final task reward of student-generated and teacher-proposed actions in a common return space for step-level comparison. ERSR analysis reveals an outcome-dependent asymmetry: student actions are more beneficial than teacher replacements on successful trajectories, whereas teacher replacements become more beneficial on failed trajectories. We further show that student answer-probe gains track student-step ERSR utility and distinguish beneficial from harmful reasoning steps. Based on these findings, we propose \textbf{Return-Referenced On-Policy Learning (R$^2$OPL)}, which reinforces student reasoning on successful trajectories and distills teacher signals on failed ones, while using group success rate for difficulty scaling and student-probe gains for step-level modulation. Experiments across reasoning benchmarks and teacher--student configurations show that R$^2$OPL consistently outperforms strong baselines. ERSR training dynamics further show that R$^2$OPL jointly exploits substantial utility from both reward- and teacher-side signals, whereas existing hybrids often leave substantial residual utility in one branch.

Comments32 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑