arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05198cs.AI

On-Policy蒸馏中关键因素是什么?数据效率与数据选择的视角

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过实证研究发现,On-Policy蒸馏中困难样本及更长思维链是性能提升的关键,仅用8个选中的困难样本训练的模型性能可匹配17K数据集基线。

中文摘要 AI 辅助

On-Policy蒸馏(OPD)已成为提升大语言模型推理能力的广泛采用的后训练范式,但OPPD中以数据为中心的机制仍相对未被充分探索。本文针对OPD中的数据效率与数据选择开展实证研究:首先探究极端场景——仅用1个样本训练OPD(即1-shot OPD),结果发现1-shot OPD在所有采样训练样本中均持续有效,且更难的样本通常能带来更优的性能提升;接着探究学生模型在训练数据中性能提升的实际驱动因素,分析表明该提升并非由高token熵驱动,而是由困难问题自然产生的更长的思维链(CoT)路径驱动,在更长的CoT上训练有助于在长推理过程中保持与教师模型更紧密的对齐,并学习短CoT中通常缺失的关键思维模式,如反思(例如“Alternatively”)。基于这些见解,本文提出一种简单的数据选择方法,仅选择困难样本用于训练,即使是完全超出教师模型能力的“不可解”样本也能成功使用。在4个参数规模从15亿到70亿的模型上开展的实验显示,仅用8个选中的困难样本训练学生模型,其性能就能达到使用17K数据集作为基线的水平。

英文摘要

On-Policy Distillation (OPD) has emerged as a widely adopted post-training paradigm for enhancing large language models in reasoning domains. However, the data-centric mechanisms in OPD remain relatively underexplored. This paper presents an empirical study of data efficiency and data selection in OPD. We begin by investigating an extreme setting: training OPD on only one example, namely 1-shot OPD. Surprisingly, we find that 1-shot OPD is consistently effective across all sampled training examples and harder examples often yield superior performance gain. We next investigate what actually drives the student model's improvement in the training data. Our analysis reveals that the improvement is not driven by high token entropy, but the longer CoT paths which hard problem naturally generate. Training on longer CoT can help maintain closer alignment with the teacher over a long reasoning horizon, and learn critical thinking patterns usually missing in short CoTs, such as reflection (e.g., "Alternatively"). Based on these insights, we propose a simple data selection method that selects only hard examples for training, where even "unsolvable" examples that completely exceed the teacher's capability can be successfully used. Our experiments conducted on five models ranging from 1.5B to 8B show that training the student model on only 8 selected hard examples matches the performance of the 17K dataset baseline. Code and models will be publicly available.

发表机构

  • Tsinghua University(清华大学)
  • Meituan(美团)

机构由 AI 辅助整理,请以论文原文为准。

↑