arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从学生状态下的教师续写中学习

Learning from Teacher Continuations at Student States

Haojin Wang, Dylan Zhang, Huaibo Chen, Suhao Yu, Yihang Sun, Zhanyang Jin, Jiaying Ye, Dianqi Li, Prasanna Sattigeri, Kamal Youcef-Toumi, Hao Peng

arXiv 2609.36246首次发表:更新:

发表机构

University of Illinois at Urbana-Champaign; Massachusetts Institute of Technology; University of Pennsylvania; University of Washington; International Business Machines(伊利诺伊大学厄巴纳-香槟分校; 麻省理工学院; 宾夕法尼亚大学; 华盛顿大学; 国际商业机器公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OLIVE通过学生生成前缀、教师续写并更新学生,解决蒸馏中的协变量偏移、碎片监督和概率访问问题,在推理和智能体任务上优于现有方法,并显著减少训练时间。

AI 中文摘要

我们提出OLIVE(在线干预)。在每次迭代中,不断演化的学生策略生成一个新的前缀,教师自回归地续写该前缀,学生则通过基于教师生成token的交叉熵进行更新。每个设计选择都针对现有蒸馏方法的相应局限性:(1)在固定教师轨迹上的离线监督微调(SFT)中存在的顺序协变量偏移,(2)在token级在线策略蒸馏(OPD)中前缀失败导致的碎片化监督,以及(3)在分布匹配蒸馏中需要访问教师token概率的问题。OLIVE在可比GPU小时成本下,比OPD(采用top-16 KL近似)实现了更高的推理性能。我们的异步实现进一步将OLIVE的总训练时间减少了23.8%。我们在硬推理任务和智能体任务上评估OLIVE,这些任务反映了现代后训练场景,在相同训练预算下,OLIVE始终优于现有蒸馏方法。通过从演化中的学生重新生成前缀,OLIVE在离线蒸馏平台期后继续改进,同时更好地保持学生的通用能力和可塑性。仅使用GPT-5.4-mini的文本,持续使用OLIVE训练比从同一教师进行离线SFT在ScienceWorld上高出13%。这些结果支持OLIVE作为一种有效且高效的在线语言模型蒸馏方法。

英文摘要

We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution-matching distillation. OLIVE achieves higher reasoning performance than OPD (with a top-16 KL approximation) at comparable GPU-hour cost. Our asynchronous implementation further reduces OLIVE's total training time by 23.8\%. We evaluate OLIVE on both hard reasoning tasks and agentic tasks which reflects modern post-training scenarios, and it consistently outperforms existing distillation methods under the same training budget. By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student. Using only text from GPT-5.4-mini, continuously training with OLIVE outperforms offline SFT from the same teacher by 13\% on ScienceWorld. These results support OLIVE as an effective and efficient approach to online language-model distillation.

CommentsHW and DZ contributed equally and share the first-authorship. Dylan Zhang is project lead

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑