arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18321cs.LGcs.AI

轨迹可学习性:面向不完美教师的离线在线策略蒸馏

Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers

Yihao Ai, Weilong Yan

首次发表
浏览论文内容

中文总结 AI 辅助

针对不完美教师监督的离线在线策略蒸馏,提出基于教师成功问题定义可学习性信号并加权损失的方法,无需额外生成,在数学推理和代码生成上提升最多2.7个百分点,且计算成本更低。

中文摘要 AI 辅助

离线在线策略蒸馏(Offline On-Policy Distillation)通过一次性收集学生轨迹和教师监督,并在整个优化过程中重复使用它们来提高效率。这种重复使用使得不完美的监督变得持续存在。由于即使是强大的教师也可能失败,我们提出一个问题:\u201c从不完美的教师监督中,什么仍然是可以学习的?\u201d教师失败只是一个粗略的问题级信号,并不表示沿着相关学生轨迹的所有监督都是无用的。一个自然的替代方案是估计教师沿轨迹的可恢复性,但重复的延续在很大程度上抹去了离线蒸馏的效率优势。我们转而使用教师成功解决的问题来定义一个廉价的参考,以衡量学生可以学习什么。我们在教师成功的问题上进行训练,并衡量来自教师失败问题的轨迹中每个观察到的词元的似然如何变化。我们将这些带符号的似然变化用作操作性的\u201c可学习性信号\u201d:更大的增加表示行为被仅成功学习更强烈地促进。我们将此信号聚合成原始蒸馏损失的轨迹级权重。与基于延续的估计不同,我们的可学习性不需要额外的生成,并且可以从存储的轨迹和模型检查点中一次性计算。在数学推理和代码生成方面,我们的方法将离线OPD基线提高了最多2.7个百分点,并在多个基准上匹配或优于在线OPD变体。尽管增加了仅成功的蒸馏阶段,它使用2个GPU和约22个GPU小时,而代表性的在线OPD方法使用3个GPU和36-48个GPU小时。

英文摘要

Offline on-policy distillation gains efficiency by collecting student trajectories and teacher supervision once and reusing them throughout optimization. The same reuse makes imperfect supervision persistent. Since even strong teachers can fail, we ask \emph{what remains learnable from imperfect teacher supervision?} Teacher failure is only a coarse problem-level signal and does not imply that all supervision along the associated student trajectory is unhelpful. A natural alternative is to estimate teacher recoverability along the trajectory, but repeated continuations largely erase the efficiency advantage of offline distillation. We instead use teacher-successful problems to define a cheap reference for what the student can learn. We train on teacher-successful problems and measure how the likelihood of each observed token in trajectories from teacher-failed problems changes. We use these signed likelihood changes as an operational \emph{learnability signal}: larger increases indicate behavior more strongly promoted by successful-only learning. We aggregate this signal into trajectory-level weights for the original distillation loss. Unlike continuation-based estimates, our learnability requires no additional generation and can be computed once from stored trajectories and model checkpoints. Across mathematical reasoning and code generation, our method improves an offline OPD baseline by up to 2.7 percentage points and matches or outperforms online OPD variants on multiple benchmarks. Despite the additional successful-only distillation stage, it uses 2 GPUs and about 22 GPU hours, compared with 3 GPUs and 36--48 GPU hours for representative online OPD methods.

发表机构

  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑