arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11247cs.LGcs.CL

为什么在线策略蒸馏有时会失败:学习信号消失

Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals

Lei Zhao, Qichao Zhao, Bowen Zuo, Qishi Zhan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究探究在线策略蒸馏(OPD)失败的机制,发现大规模教师模型的OPD会出现早期损失平台,推测有限的表示适应可能导致学习信号崩溃,为理解OPD的性能差异提供了条件性解释。

中文摘要 AI 辅助

在线策略蒸馏(OPD)可实现语言模型间的有效能力迁移,但其失败的机制尚未完全明晰。在代码生成与数学推理任务中,采用更大规模教师模型的OPD会出现早期损失平台,经200次更新后平均最终损失降低25.1%,而初始学生模型经进一步强化学习(RL)训练得到的自RL教师模型,该数值为96.2%。为理解这一差异,我们在小学习率极限下将OPD分析为理想化的连续时间动力系统。我们的训练日志诊断将这些平台与基于梯度的学习信号代理的早期下降关联起来,尽管仍存在大量损失;这些测量并未确定潜在梯度减弱的原因。我们进一步证明,在正则条件下,当教师模型与初始学生模型在共享参数化中足够接近时,存在局部恢复保证,为实验中自RL教师模型的成功提供了条件性解释。在有和无损失平台的运行中,我们观察到相对参数变化较小(0.025%-0.098%),且学生模型在OPD前后的表示具有高度相似性(各层线性CKA均>0.98)。这些观察表明,有限的表示适应可能导致学习信号崩溃,这一假设仍有待验证。代码可在该https URL获取。

英文摘要

On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood. Across code generation and mathematical reasoning, OPD with larger-scale teachers exhibits early loss plateaus, with an average final loss reduction of 25.1% after 200 updates, compared with 96.2% for self-RL teachers, obtained by further reinforcement learning (RL) training of the initial student. To understand this difference, we analyze OPD as an idealized continuous-time dynamical system in the small-learning-rate limit. Our training-log diagnostics associate these plateaus with an early decline in a gradient-based learning-signal proxy while substantial loss remains; these measurements do not establish why the underlying gradient weakens. We further prove a local recovery guarantee for teachers sufficiently close to the initial student in a shared parameterization under regularity conditions, offering a conditional explanation for the success of self-RL teachers in our experiments. Across runs with and without loss plateaus, we observe small relative parameter changes (0.025-0.098%) and high similarity between the student's representations before and after OPD (linear CKA $>0.98$ across layers). These observations suggest that limited representation adaptation may contribute to learning-signal collapse, a hypothesis that remains to be tested. Code is available at https://github.com/leizhao7/opd-learning-signals.

发表机构

  • University of Pennsylvania(宾夕法尼亚大学)
  • Tsinghua University(清华大学)
  • University of California, Riverside(加州大学河滨分校)
  • Marquette University(马凯特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑