arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24987cs.LGcs.AI

D³-MOPD:面向高效多教师蒸馏的自适应动态域调度

D$^3$-MOPD: Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng, Guojian Zhan, Wenliang Chen, Yao Hu, Min Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多教师在线蒸馏中固定域混合比例的缺陷,提出零开销调度器 D³-MOPD,通过跟踪 KL 轨迹自适应调整域采样比例,在 Qwen3.6-35B-A3B 模型上显著提升蒸馏效率与性能。

中文摘要 AI 辅助

多教师在线蒸馏(MOPD)通过最小化学生模型自身 rollout 上的各域反向 KL 散度,将多个领域专家教师模型蒸馏为单个学生模型。现有方法通常在训练前固定各域的数据混合比例,却忽略了不同域的收敛速率差异巨大:部分域会较早进入平台期,而其他域则会在整个训练预算内持续改进。固定混合比例因此会在快速收敛的域上浪费计算资源,同时导致较慢收敛的域训练不足。为解决该问题,我们提出 D³-MOPD(面向 MOPD 的动态域调度),这是一种零开销调度器,它复用训练过程中已产生的各域反向 KL 信号,在线自适应调整域混合比例。该调度器作为训练流程外的异步进程运行,一个外部监控器会定期跟踪每个域的 KL 轨迹,估算剩余提升空间和当前改进速率,并据此调整域采样比例,且不改变核心训练循环。我们的 D³-MOPD 可自然扩展到任意数量的域,当更多域引入更多样的收敛模式供调度器利用时,其预期收益会随之增长。在从四个领域专家教师模型蒸馏得到的 Qwen3.6-35B-A3B 学生模型上,D³-MOPD 缩小了平均学生-教师性能差距的 97%,而普通 MOPD 仅缩小 63%;在约 3 倍的 rollout 步骤减少量下达到相同的峰值性能,且在 7 个基准测试中的 3 个上超过了专家教师模型。

英文摘要

Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollouts. Existing approaches typically fix the per-domain data mixture before training, overlooking the fact that different domains converge at substantially different rates: some plateau early while others continue to improve throughout the training budget. A fixed mixture therefore wastes compute on fast-converging domains and undertrains slower-converging ones. To address this, we propose D$^3$-MOPD (Dynamic Domain ScheDuling for MOPD), a zero-overhead scheduler that repurposes the per-domain reverse-KL signal already produced during training to adapt the domain mixture online. Running asynchronously outside the training process, an off-process watcher periodically tracks each domain's KL trajectory, estimates remaining headroom and current improvement rate, and accordingly adjusts the domain sampling ratios without altering the core training loop. Our D$^3$-MOPD scales naturally to arbitrary numbers of domains, and the expected benefit grows as more domains introduce more diverse convergence patterns for the scheduler to exploit. On a Qwen3.6-35B-A3B student distilled from four domain-expert teachers, D$^3$-MOPD closes 97% of the average student-to-teacher performance gap, compared with 63% for vanilla MOPD, reaches the same peak performance with an approximately 3$\times$ reduction in rollout steps, and surpasses the specialist teachers on three of seven benchmarks.

↑