arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

残差优势:用于可验证奖励强化学习的学生相对教师指导

Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards

Xiaobing Chen, Zhiqi Pang

arXiv 2610.11519首次发表:更新:

发表机构

Harbin Engineering University; Tencent; Harbin Institute of Technology(哈尔滨工程大学; 腾讯; 哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出残差优势(RA)及CoRA方法,结合GRPO或REINFORCE++,在三个数学基准的24次比较中提升了序列优势算法的Avg@8和Pass@8,效果优于仅教师的OPD。

AI 中文摘要

可验证奖励强化学习(RLVR)和在线策略蒸馏(OPD)已成为后训练推理模型的两大主要范式。RLVR为每个响应提供单一结果标签,不对响应内部的步骤单独赋予信用;OPD在学生访问的前缀处提供token级指导,但其逐点信号无法直接反映教师与学生在整个词汇表上的分歧模式。密集、无界的对数比监督可放大教师的影响力,但当学生的解路径偏离教师时,强求解器未必是合适的指导者。我们提出残差优势(\textbf{Residual Advantage, \textbf{RA}}),将师生概率残差视为有界单步奖励,减去学生策略下对应的状态值以形成标准优势,并在每个响应内对结果进行中心化处理后,再将其添加到验证器优势中。该指导项在每个响应内均值为零,因此验证器优势仍为响应的均值标签,教师仅在响应内部的步骤间重新分配信用。\textbf{CoRA}进一步在同一已评分学生批次上用验证器优势更新教师LoRA,并在下一轮迭代的残差中使用更新后的教师,使指导适配学生的尝试。使用Qwen3-1.7B-Base和Qwen3-4B-Base作为学生,Qwen3-8B作为教师,\textbf{RA}与GRPO或REINFORCE++结合后,在三个数学基准的全部24次比较中均优于基础序列优势算法,将宏观Avg@8提升1.7至3.6个百分点,Pass@8提升3.9至6.3个百分点;两种组合均优于仅使用教师的OPD,而\textbf{CoRA}额外提升了1.0至1.5个Avg@8百分点。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have become two main paradigms for post-training reasoning models. RLVR gives each response a single outcome label, leaving the steps inside it without separate credit. OPD provides token-level guidance at student-visited prefixes, but its pointwise signal does not directly reflect the pattern of teacher--student disagreement across the vocabulary. Dense, unbounded log-ratio supervision can amplify the teacher's influence, yet a strong solver is not necessarily a suitable guide when the student's solution paths depart from the teacher's. We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage. The guidance term has zero mean within each response, so the verifier advantage remains the response's mean label and the teacher only redistributes credit among the steps within it. \CoRA{} further updates a teacher LoRA with verifier advantages on the same scored student batch and uses the updated teacher in the next iteration's residual, adapting guidance to the student's attempts. With Qwen3-1.7B-Base and Qwen3-4B-Base students and a Qwen3-8B teacher, \RA{} combined with GRPO or REINFORCE++ improves the underlying sequence-advantage algorithm in all 24 comparisons on three mathematical benchmarks, raising macro Avg@8 by 1.7--3.6 points and Pass@8 by 3.9--6.3 points. Both combinations surpass teacher-only OPD, and \CoRA{} adds a further 1.0--1.5 Avg@8 points.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑