arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

教师应如何准备?基于学生诱导状态的强化学习用于在线策略蒸馏

How Should Teachers Be Prepared? RL on Student-Induced States for On-Policy Distillation

Xiaoyu Ma, Haoyue Liu, Zhichao Wang, Jionghao Zhu, Xiaoying Tang

arXiv 2610.04950首次发表:更新:

发表机构

School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen; Shenzhen Future Network of Intelligence Institute (FNiI-Shenzhen); Guangdong Provincial Key Laboratory of Future Networks of Intelligence, CUHK(SZ)(香港中文大学(深圳)理工学院; 深圳市未来智能网络研究院; 广东省未来智能网络重点实验室(香港中文大学(深圳)))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Prep-OPD,在蒸馏前用强化学习训练教师适应学生推理状态并纠错,显著提升数学推理蒸馏性能。

AI 中文摘要

在线策略蒸馏(OPD)通过在学生生成的轨迹上提供词级教师监督,提升小型语言模型的推理能力。然而,擅长独立解决问题的教师能否同样有效地引导学生推理?先前研究表明,当学生前缀的推理路径与教师自身不同或包含错误时,教师从这些前缀继续推理的准确性可能低于独立解决问题。为此,我们提出Prep-OPD,在蒸馏前使用强化学习(RL)训练教师适应学生现有的推理状态,并在出现错误时纠正方向。训练以最终答案正确性为奖励,优化教师从固定学生前缀开始的延续。准备好的教师随后通过轨迹引导和词级监督训练学生。我们在八个数学推理基准上评估Prep-OPD,使用Qwen3-4B-Instruct-2507作为教师,Qwen3-0.6B和Qwen3-1.7B作为学生。对于4B教师和1.7B学生,Prep-OPD相比标准OPD和最强基线Relay-OPD,平均准确率分别提升8.28和2.30个百分点。受控实验进一步表明,基于学生生成前缀的教师RL在Qwen3-1.7B上比从问题开始(有无交接)的教师RL产生更高的学生准确率。复用同一准备好的教师也提升了Qwen3-0.6B的性能。

英文摘要

On-policy distillation (OPD) improves the reasoning capabilities of small language models through token-level teacher supervision on student-generated trajectories. Yet can teachers that excel at solving problems independently also guide student reasoning effectively? Prior work shows that when student prefixes follow reasoning paths that differ from the teacher's own or contain errors, teachers can be less accurate when continuing from these prefixes than when solving problems independently. To this end, we propose Prep-OPD, which uses reinforcement learning (RL) before distillation to train the teacher to adapt to the student's existing reasoning state and correct course when errors arise. Training optimizes teacher continuations from fixed student prefixes using final-answer correctness as the reward. The prepared teacher then trains the student through trajectory guidance and token-level supervision. We evaluate Prep-OPD on eight mathematical reasoning benchmarks, using Qwen3-4B-Instruct-2507 as the teacher and Qwen3-0.6B and Qwen3-1.7B as students. With the 4B teacher and 1.7B student, Prep-OPD improves average accuracy over standard OPD and the strongest baseline, Relay-OPD, by 8.28 and 2.30 percentage points, respectively. Controlled experiments further show that teacher RL conditioned on student-generated prefixes yields higher student accuracy than problem-start teacher RL with and without handoff on Qwen3-1.7B. Reusing the same prepared teacher also improves Qwen3-0.6B.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑