发表机构
University of Chinese Academy of Sciences; Tsinghua University; Northeastern University; University of Illinois Urbana-Champaign; Johns Hopkins University(中国科学院大学; 清华大学; 东北大学; 伊利诺伊大学厄巴纳-香槟分校; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过单查询训练探究在线策略蒸馏(OPD)的训练数据作用,发现单查询 OPD 可恢复大部分全数据增益,16 个语义查询可匹配全数据效果,指出 OPD 数据过剩但算法不足,为后续研究指明方向。
AI 中文摘要
在线策略蒸馏(OPD)结合学生模型生成的rollouts(滚动输出)与教师模型提供的密集 token 级监督。现有研究主要关注其算法行为,却未明确训练数据的作用。我们在数据最小极限下,仅用单个查询进行训练以探究该作用。单次 OPD 会持续优化数百步,在各任务领域及模型族中恢复了全数据 OPD 的大部分增益。我们通过训练期间访问的状态以及学生模型与教师模型对齐的速率来解释这一结果。我们测量「状态覆盖率」,即全数据 OPD 访问的状态中,某查询集的 rollouts 能达到的比例。单个查询已能达到 71.5%,且大部分在前 100 步内完成。添加语义不同的查询会同步提升覆盖率与验证准确率,直到 16 个查询达到 98.9%,与全数据训练效果相当。然而,无论 OPD 用单个查询还是全数据集训练,对齐速率的下降幅度相似,即使是固定的状态集也需要数百步才能被吸收。因此,OPD 是数据过剩但算法不足的;其 rollouts 能快速暴露广泛的监督信号,而学生模型吸收该监督信号的速度却越来越慢。状态覆盖率结果可推广至多教师 OPD(MOPD),每个领域 16 个语义多样的查询即可达到全数据 MOPD 的效果。进一步的压力测试显示,内容稀疏的模板和域外 WildChat 查询也接近真实查询的基准。因此,任务内容与诱导的状态覆盖率可能相互分离。我们希望这些发现能引导未来研究关注 OPD 的步骤效率,并促使重新审视其近期在前沿后训练中成功背后的数据与机制。
英文摘要
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure \emph{state coverage}, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \(71.5\%\), most of it within the first 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach \(98.9\%\) and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or the whole dataset, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. Task content and induced state coverage can therefore come apart. We hope these findings direct future work toward the step efficiency of OPD, and prompt a re-examination of the data and the mechanisms behind its recent successes in frontier post-training.
Comments29 pages, 20 figures