arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

教师是方向而非终点:在在线策略蒸馏中外推RL诱导的表示残差

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

Hao Li, MeiJia Chen, Weijie Ren, Donghan Li, Zijun Tian, Jingchun Huang, Naibo Wang

arXiv 2609.36484首次发表:更新:

发表机构

University of Science and Technology of China; Rutgers University; Zhejiang University(中国科学技术大学; 罗格斯大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对在线策略蒸馏中学生难以超越教师的问题,提出RIDE方法,在表示空间直接外推RL诱导的残差方向,在多种规模与架构上稳定超越教师。

AI 中文摘要

在线策略蒸馏(OPD)训练学生模型在自身轨迹上匹配教师的下一词元分布,并已取得了显著的实证收益。广义变体允许学生通过外推输出空间中的隐式奖励来超越教师。然而,语言模型头会各向异性地衰减这种变化:教师隐藏状态中编码的大部分变化仅以一小部分权重到达对数几率,而输出空间外推所依赖的采样词元对数概率比会注入噪声,外推过程会放大这些噪声,导致训练不稳定。我们观察到,强化学习(RL)会使模型内部表示相对于其基础检查点发生偏移,并且这种偏移的方向可以在每一层进行测量。基于这一观察,我们提出了RIDE(RL诱导方向外推),直接在表示空间中外推RL诱导的变化:在每一层和每个词元位置,RIDE计算教师与其RL前检查点之间的残差,并将学生的隐藏状态回归到沿此残差方向超越教师的目标。在采样轨迹的条件下,该回归等价于在教师为中心的二次惩罚下最大化由残差定义的线性方向奖励,这明确说明了目标如何沿RL诱导方向移动学生,同时限制其与教师的偏差。在跨越不同规模、架构和预训练谱系的四组基础/RL教师对上,RIDE在每一对上都接近或超过了RL训练的教师,并且是唯一在均值上做到这一点的方法,同时它始终优于输出空间外推,后者在教师接近其基础模型时会降低学生性能。项目页面:此https URL。

英文摘要

On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however, attenuates this change anisotropically: much of the change encoded in the teacher's hidden states reaches the logits at a small fraction of its weight, and the sampled-token log-probability ratios on which output-space extrapolation relies inject noise that the extrapolation amplifies, making training unstable. We observe that reinforcement learning (RL) shifts a model's internal representations relative to its base checkpoint, and that the direction of this shift can be measured at every layer. Motivated by this observation, we propose RIDE (RL-Induced Direction Extrapolation), which extrapolates the RL-induced change directly in representation space: at every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student's hidden states toward targets displaced beyond the teacher along this residual. Conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward defined by the residual under a quadratic penalty centered at the teacher, which makes explicit how the objective moves the student along the RL-induced direction while limiting its deviation from the teacher. Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base. Project page: https://github.com/xixixixixxxx/RIDE.

Comments19 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑