arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无数据在线策略蒸馏

Data-Free On-Policy Distillation: How Far Can We Go Without External Data?

Gengsheng Li, Mao Zheng, Mingyang Song, Jie Sun, Zeyuan Liu, Ruiqi Liu, Tianyu Yang, Qiyong Zhong, Haiyun Guo, Junfeng Fang, Shiming Xiang, Jinqiao Wang, Tat-Seng Chua

arXiv 2609.14193首次发表:更新:

发表机构

School of Artificial Intelligence, University of Chinese Academy of Sciences; Foundation Model Department, Tencent; National University of Singapore; Wuhan AI Research(中国科学院大学人工智能学院; 腾讯基础模型部; 新加坡国立大学; 武汉人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出无数据在线策略蒸馏(DF-OPD),发现OPD对数据不敏感,教师自生成问题即可匹配甚至超越真实数据,在多教师蒸馏中填补98.5%的提升空间。

AI 中文摘要

在线策略蒸馏(OPD)已成为前沿后训练流程的标准组成部分,然而其训练数据实际贡献了多少在很大程度上仍未得到检验。在实践中最常见的两种师生配对中,我们发现OPD对其数据几乎不敏感:8个提示词即可媲美一个包含17k问题的数据集,而三个独立构建的、难度和师生KL散度相差数倍的数据集产生的训练曲线几乎无法区分。造成这一现象的原因有两个。首先,OPD中数据的单位是提示词所导致的状态,而非提示词本身:随着采样的持续,单个提示词会不断暴露新的教师修正,而额外提示词的边际价值在八个之后便趋于崩溃。其次,用竞技编程替代数学仍能恢复超过百分之九十的领域内收益,这表明OPD传递的是教师的推理模式,而非与数据相关的知识。我们将这一发现推向极致,提出了无数据在线策略蒸馏(DF-OPD),其中教师在简单提示词下自行编写训练问题——无需外部数据,无需过滤——仅留下一个由两个策略组成的系统。DF-OPD匹配甚至超越了真实数据,并且其产生的问题在训练动态的三个关键诊断指标上追踪了教师自身的后训练数据,而其他真实数据集则无法做到。应用于多教师蒸馏时,通常需要从往往难以获取的后训练数据中推导(提示词,领域)对,而1k个自生成问题即可填补98.5%的可用提升空间,甚至超越了使用7k个真实样本所达到的96.6%。此外,这些结果共同促使我们重新评估数据在OPD中所扮演的角色。

英文摘要

On-policy distillation (OPD) is increasingly applied to frontier foundation model post-training. Prior work in this area has largely focused on algorithmic advances, yet it remains unclear how much OPD depends on its training questions and, in particular, how far this dependence can be reduced. Across two representative single-teacher OPD settings, we find that training on 8 real prompts yields performance comparable to training on 17k problems, while datasets differing substantially in measured difficulty and initial distillation gap yield similar outcomes. Our analyses suggest two complementary explanations: repeated sampling could allow even a few prompts to expose substantial teacher supervision, while OPD transfers generalizable reasoning capabilities beyond dataset-specific knowledge. Building on these observations, we next propose a data-free on-policy distillation (DF-OPD) setting to investigate whether the system can supply the training questions itself, eliminating the need for external data. With 64 self-generated questions obtained without seed examples, DF-OPD yields performance comparable to full-data OPD in both single-teacher settings. This finding also holds in multi-teacher OPD: across mathematics, code, and instruction following, 1k generated questions achieve performance comparable to training on approximately 7k real post-training examples. We further explore whether OPD can operate even without explicit training questions. The experiments show that this is effective only in limited cases, where the student unexpectedly generates and answers its own questions, thereby reducing the process to an implicit form of DF-OPD. Together, these findings invite a reassessment of the role of training data in on-policy distillation. Code is available at https://github.com/Ryuki661/DF-OPD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑