发表机构
University of Science and Technology of China; Peking University; IQuest Research; MBZUAI; Zhejiang University(中国科学技术大学; 北京大学; IQuest研究院; 穆罕默德·本·扎耶德人工智能大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究大型语言模型策略内蒸馏(OPD)的泛化特性,发现其迁移教师推理行为而非答案,同源配对泛化性强,异源配对适配性有限,多教师组合存在能力跷跷板效应,为诊断多教师OPD提供了视角。
AI 中文摘要
策略内蒸馏(OPD)通过监督学生自身策略采样的轨迹来迁移教师的能力,但其泛化行为仍未得到充分理解,因为多数研究仅在单一领域及接近训练数据的基准上评估OPD。我们开展了一项控制变量研究,每次改变一个泛化因素,涵盖域内分布偏移、跨域迁移及多教师设置。我们发现OPD迁移的是教师的推理行为而非特定问题的答案:训练难度几乎不产生影响,甚至教师从未解决的问题也有用。迁移高度依赖教师与学生的起源关系:同源配对能让学生在语言、推理范围甚至其他领域接近教师,而异源配对大多仅适配训练分布。这种广泛影响是一把双刃剑:由于将提示路由到领域专家无法限制每位教师的影响,组合它们会在其能力间产生依赖混合的跷跷板效应。这些结果阐明了OPD泛化的适用场景,并为诊断多教师OPD提供了有用视角。
英文摘要
On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.
CommentsUnder Review