先行训练,回传蒸馏:大型语言模型的在策略自蒸馏引导方法
Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models
- ShanghaiTech University(上海科技大学)
- BIGAI(北京通用人工智能研究院)
- Hefei University of Technology(合肥工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出B-OPSD方法,通过临时超前训练策略获得未来教师,再回传蒸馏监督,提升大型语言模型在策略自蒸馏效果,在数学推理任务上显著优于标准OPSD。
AI中文摘要:
在策略自蒸馏(OPSD)通过让具有特权信息的自我教师模型在模型自身的轨迹上提供密集的令牌级监督,从而改进大型语言模型。然而,现有方法通常从当前、初始或缓慢平均的策略状态构建自我教师,这使得监督质量受限于教师利用特权信息的能力。我们探究模型自身的优化进展是否可以被回收利用,以构建更强的自我教师。在本文中,我们引入了引导式在策略自蒸馏(B-OPSD),该方法临时将策略训练至未来状态以获得未来教师,将学生恢复至原始策略状态,然后利用未来教师监督重新开始的学生。未来教师通过两种互补方式改进监督:它能生成更可靠的特权轨迹,并基于这些轨迹,在重新开始的学生在策略轨迹上提供更具信息量的令牌级目标。在Qwen3-4B和Qwen3-8B上的数学推理实验表明,在两种设置下均优于标准OPSD,包括在rollout特权设置中从27.50提升至41.30,以及从48.80提升至64.44。我们的发现指向一个更广泛的自我改进模型原则:未来的学习进展可以向后蒸馏,保留已获得的知识,同时引导超越产生该知识的优化状态。
英文摘要:
On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the quality of supervision constrained by the teacher's ability to exploit privileged information. We ask whether the model's own optimization progress can instead be recycled into a stronger self-teacher. In this paper, we introduce Bootstrapped On-Policy Self-Distillation (B-OPSD), which temporarily trains the policy ahead to obtain a future teacher, restores the student to the original policy state, and then uses the future teacher to supervise the restarted student. The future teacher improves supervision in two complementary ways, it can generate more reliable privileged trajectories and, conditioned on them, provide more informative token-level targets along the restarted student's on-policy trajectories. Experiments on mathematical reasoning with Qwen3-4B and Qwen3-8B show consistent improvements over standard OPSD in both settings, including gains from 27.50 to 41.30 and from 48.80 to 64.44 in the rollout-privileged setting. Our findings point to a broader principle for self-improving models that future learning progress can be distilled backward, preserving acquired knowledge while bootstrapping beyond the optimization state that produced it.