发表机构
University of Maryland; Microsoft Research; MBZUAI(马里兰大学; 微软研究院; 穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出W2S-OPD框架,从多个弱模型蒸馏提升强学生模型,在数学、代码基准测试中优于OPD,可让学生超越领域教师,监督源更弱时仍能提升性能。
AI 中文摘要
在线策略蒸馏(OPD)是一种通过让学生模型对齐教师模型在自身 rollout 上的 token 级分布来实现大语言模型(LLM)能力迁移的有效范式。现有主流方法假设教师模型的能力至少不弱于学生模型:要么将更大模型蒸馏为更小模型,这在不存在更大教师的前沿场景失效;要么整合多个从共享基础模型训练的领域专家,这需要在学生模型规模上进行成本高昂的训练。本文提出弱到强在线策略蒸馏(W2S-OPD),一种简单却有效的 OPD 框架,通过从多个弱模型蒸馏来提升强学生模型。W2S-OPD 从正模型和负模型的对比对在 logit 空间构建代理教师,两者均小于学生模型且易获取;它们的 logit 差值分离出能力方向,将其添加到学生模型自身的基础模型中,得到既耦合该方向又与学生模型分布相近的代理教师。学生模型随后通过最小化自身 rollout 上的逐 token 反向 KL 来蒸馏该代理教师。本文将对比对实例化为三类:i)强化学习后(post-RL)专家与强化学习前(pre-RL)初始化的对比,分离强化学习注入的技能;ii)更大基础模型与更小基础模型的对比,分离规模带来的能力;iii)带有正确提示与错误提示的小基础模型的对比,分离实例层面指向解决方案的方向。在四个数学和三个代码基准测试中,W2S-OPD 优于 OPD,使学生模型超越领域教师,且即使所有监督源都更弱时仍能持续提升学生模型。分析显示不同对比会产生不同信号:强化学习后与提示对比强调推理框架,规模对比强调求解过程。我们的代码将在该 https URL 公开。
英文摘要
On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabilities across LLMs. Prevailing approaches assume a teacher at least as capable as the student: they either distill a larger model into a smaller one, which fails at the frontier where no larger teacher exists, or consolidate multiple domain experts trained from a shared base, which requires costly training at the student's scale. We introduce Weak-to-Strong On-Policy Distillation (W2S-OPD), a simple yet effective OPD framework that improves the strong student by distilling from multiple weak models. W2S-OPD constructs a proxy teacher in logit space from a contrast pair of a positive and a negative model, both smaller than the student and cheap to obtain. Their logit difference isolates the capability direction, which is added to the student's own base model, yielding a proxy teacher that couples this direction while staying distributionally adjacent to the student. The student then distills it by minimizing the per-token reverse KL on its own rollouts. We instantiate the contrast pair as i) a post-RL expert against its pre-RL initialization, isolating the skill RL instills, ii) a larger against a smaller base model, isolating the capability from scale, and iii) a small base model with correct versus wrong hints, isolating the instance-level direction toward the solution. Across four math and three code benchmarks, W2S-OPD outperforms OPD, enables the student to surpass the domain teacher, and keeps improving the student even when every supervision source is weaker. Analysis shows different contrasts yield distinct signals: the post-RL and hint contrasts emphasize reasoning frameworks, while the scale contrast emphasizes the solving procedure. Our code will be available at https://github.com/Yu-Fangxu/W2S-OPD.
CommentsTechnical Report