arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种症状,三个杠杆:对在线策略自蒸馏的批判性综述

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

Justin Robert, Raheel Qader

arXiv 2608.25936首次发表:更新:

发表机构

OVHai LLM(OVHai大模型)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述针对在线策略自蒸馏(OPSD)存在的推理路径崩溃问题,从信号加权、特权信息性质、指导衰减三个杠杆展开分析,以数学推理为范围,提供统一术语并区分已确定与争议内容。

AI 中文摘要

在线策略蒸馏(On-policy distillation)在语言模型自身生成的内容上对其进行训练,同时由教师模型逐词对这些内容打分,它将模仿学习的密集监督与强化学习的在线策略采样相结合,但需要一个更大的模型作为教师模型。在线策略自蒸馏(On-Policy Self-Distillation, OPSD)消除了这一成本,教师模型是模型本身,以学生模型在测试时不会拥有的特权信息为条件,例如参考解决方案、规划结果或环境反馈,教师模型并不比学生模型更强,只是信息更充分。早期结果令人鼓舞,其准确率可与强化学习相媲美,但生成的 token 仅为后者的一小部分。不过,产生信号的相同不对称性也会使其产生偏差,目前该领域主要存在一种失败模式:崩溃(collapse),即模型可生成的推理路径集合逐渐变窄。崩溃并非 OPSD 所特有,尽管特权信息会加剧这一问题。本综述将崩溃视为一种由三个杠杆控制的症状:(i)信号应用的位置,即如何对 token 进行加权;(ii)教师模型所看到的内容,即特权信息的性质;(iii)信号变化的时间,即教师模型的动态性和指导的衰减。我们将范围限制在该方法起源且失败模式记录最充分的数学推理领域,未报告新的实验,贡献在于结构层面:为不同论文中命名各异的现象提供了统一术语,并明确区分了已确定的内容与仍有争议的内容。

英文摘要

On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were promising, with accuracy comparable to reinforcement learning at a fraction of the generated tokens. But the same asymmetry that produces the signal also biases it. One failure mode now dominates the field: collapse, the progressive narrowing of the set of reasoning paths the model can produce. Collapse is not specific to OPSD, though privileged information aggravates it. This review treats collapse as a symptom governed by three levers: (i) where the signal is applied, that is, how tokens are weighted; (ii) what the teacher is shown, that is, the nature of the privileged information; and (iii) when the signal changes, that is, the teacher's dynamics and the decay of guidance. We restrict our scope to mathematical reasoning, where the method originated and where its failure modes are best documented. We report no new experiments. The contribution is structural: a shared vocabulary for phenomena named differently across papers, and a clear line between what is settled and what is still disputed.

Comments30 pages, 4 figures. Survey / critical review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑