Open-MOPD:诊断与修复多教师在线蒸馏中的能力不平衡
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
- SIA-Lab of Tsinghua AIR and ByteDance Seed(清华大学AIR与字节跳动种子SIA实验室)
- Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
- Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对多教师在线蒸馏(M-OPD)的能力不平衡问题,提出 Open-MOPD 框架,通过三种机制将提升空间恢复率从 35.6% 提升至 83.4%,并开源了相关方案与评估套件。
AI中文摘要:
多教师在线蒸馏(M-OPD)作为一种有前景的范式,通过密集的 token 级奖励监督,将领域专业化的强化学习(RL)专家整合为单一的通用学生模型。尽管其实际应用取得了成功,但多教师能力整合的优化动力学仍未被充分理解,且明显缺乏公开、严格可复现的方案。在本研究中,我们基于 SmolLM3-3B-Base 建立了带有神经常规路由的受控 M-OPD 基准,将能力整合与路由歧义隔离开来。我们的研究揭示了显著的能力整合差距:标准 M-OPD 相对于领域路由神经常规集成仅捕获了 35.6% 的可用提升空间,且指令遵循等简洁任务出现严重退化和过早停滞。至关重要的是,我们表明这种失败并非源于梯度冲突,而是源于 token 级优化预算的严重错配。这种问题由三个正交因素驱动:各领域间的结构性序列长度差异、因学习率不均匀导致的动态收敛漂移,以及异步策略更新带来的多步奖励过时。为解决这些不平衡,我们引入了 Open-MOPD,这是一个结合了 token 份额平衡、感知差距的动态预算分配和学生奖励刷新的原则性框架。这些机制共同系统地恢复了跨领域平衡,将单一可部署学生模型的提升空间恢复率从 35.6% 提升至 83.4%。我们在学术可及的硬件预算上完全开源了端到端的后训练方案、训练轨迹和评估套件。
英文摘要:
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision. Despite its practical success, the optimization dynamics governing multi-teacher capability integration remain poorly understood, and open, rigorously reproducible recipes are conspicuously lacking. In this work, we establish a controlled M-OPD benchmark on SmolLM3-3B-Base with oracle routing, isolating capability integration from routing ambiguity. Our investigation reveals a pronounced capability integration gap: standard M-OPD captures only 35.6% of the available headroom relative to a domain-routed oracle ensemble, with concise tasks such as instruction following suffering severe degradation and premature stagnation. Crucially, we show that this failure stems not from gradient conflict, but from a severe misallocation of the token-level optimization budget. This pathology is driven by three orthogonal factors: structural sequence-length disparities across domains, dynamic convergence drift due to non-uniform learning rates, and multi-step reward staleness from asynchronous policy updates. To resolve these imbalances, we introduce Open-MOPD, a principled framework incorporating token-share balancing, gap-aware dynamic budget allocation, and student reward refresh. Together, these mechanisms systematically restore cross-domain balance, elevating headroom recovery from 35.6% to 83.4% in a single deployable student. We fully open-source our end-to-end post-training recipe, training trajectories, and evaluation suites on an academically accessible hardware budget.