arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

组合每位教师所学:通过教师相对偏移实现多教师在线策略蒸馏

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

Hejian Sang, Zhengze Zhou, Shayan Mohajer Hamidi, Xiaomin Li, Rohit Jain, Alborz Geramifard

arXiv 2610.10460首次发表:更新:

发表机构

Iowa State University; LinkedIn; Harvard University(爱荷华州立大学; 领英; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Δ-MOPD,通过转移教师相对基础模型的偏移并重新锚定到学生初始化,改善多教师在线策略蒸馏,在组合和路由设置中提升性能并减少顺序偏差。

AI 中文摘要

多教师在线策略蒸馏(MOPD)用于两种设置。在共同领域组合中,多个教师对来自同一提示域的每个学生轨迹进行评分,其信号构成单一目标;在路由领域蒸馏中,来自不同领域的提示被分配给相应的专家。两种设置通常转移每位教师的最终策略,这混合了训练后改变的内容与从教师基础模型继承的偏好。我们引入$\Delta$-MOPD,该方法转移每位教师的教师减基础对数偏移,并以学生的冻结初始化重新锚定,并在保持教师选择固定的情况下,在两种设置中将其与最终监督进行比较。我们首先揭示了阻碍最终转移的机制:继承的基础拉动力可能超过训练后偏移。移除它降低了教师项范数比率和目标-学生KL散度。在我们的实验中,结果表明当教师信号在状态上组合时,偏移目标特别有用。使用三个组合教师时,$\Delta$-MOPD在数学和五个基准点上分别超过最终组合$4.11$和$1.95$分;使用两个时,其精度与最终组合相当。在分阶段路由下,它在两种阶段顺序中均实现更高的平均性能,并将观察到的顺序差距从$10.50$分降至$6.42$分。在交错路由下,每次更新涉及一个教师,两种目标表现相当。分阶段结果提供了支持性证据,表明该益处可能扩展到跨训练阶段累积的信号。因此,目标构建是MOPD中独立的设计轴,与教师选择互补。

英文摘要

Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both settings usually transfer each teacher's endpoint policy, which mixes what post-training changed with preferences inherited from the teacher's base. We introduce $Δ$-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed. We first expose the mechanism that impedes endpoint transfer: inherited base pull can exceed the post-training shift. Removing it reduces the teacher-term norm ratio and target--student KL. Across our experiments, the results suggest that shift targets are particularly useful when teacher signals are combined at a state. With three composed teachers, $Δ$-MOPD exceeds endpoint composition by $4.11$ Math and $1.95$ five-benchmark points; with two, it matches endpoint accuracy. Under phased routing, it achieves higher mean performance in both phase orders and reduces the observed order gap from $10.50$ to $6.42$ points. Under interleaved routing, where each update involves one teacher, the two targets perform comparably. The phased results provide supporting evidence that the benefit may extend to signals accumulated across training phases. Target construction is thus an independent design axis in MOPD, complementary to teacher selection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑