arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

硬币皆有两面:关于大型语言模型的策略内蒸馏中泛化的双重性质

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

Zhaoyi Li, Deyang Kong, Yuan Wei, Evan Yang, Ranran Shen, Mahardika Krisna Ihsani, Ming Yang, Wei Zhang, Chuan Hao, Jian Yang, Ran Tao, Bryan Dai, Shikun Zhang, Wei Ye, Ying Wei, Defu Lian

arXiv 2608.16647首次发表:更新:

发表机构

University of Science and Technology of China; Peking University; IQuest Research; MBZUAI; Zhejiang University(中国科学技术大学; 北京大学; IQuest研究院; 穆罕默德·本·扎耶德人工智能大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究大型语言模型策略内蒸馏(OPD)的泛化特性,发现其迁移教师推理行为而非答案,同源配对泛化性强,异源配对适配性有限,多教师组合存在能力跷跷板效应,为诊断多教师OPD提供了视角。

AI 中文摘要

策略内蒸馏(OPD)通过监督学生自身策略采样的轨迹来迁移教师的能力,但其泛化行为仍未得到充分理解,因为多数研究仅在单一领域及接近训练数据的基准上评估OPD。我们开展了一项控制变量研究,每次改变一个泛化因素,涵盖域内分布偏移、跨域迁移及多教师设置。我们发现OPD迁移的是教师的推理行为而非特定问题的答案:训练难度几乎不产生影响,甚至教师从未解决的问题也有用。迁移高度依赖教师与学生的起源关系:同源配对能让学生在语言、推理范围甚至其他领域接近教师,而异源配对大多仅适配训练分布。这种广泛影响是一把双刃剑:由于将提示路由到领域专家无法限制每位教师的影响,组合它们会在其能力间产生依赖混合的跷跷板效应。这些结果阐明了OPD泛化的适用场景,并为诊断多教师OPD提供了有用视角。

英文摘要

On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑