发表机构
University of Science and Technology of China; Tencent; Tsinghua University; Institute of Automation, Chinese Academy of Sciences; Zhongguancun Academy; Nankai University; National University of Singapore(中国科学技术大学; 腾讯; 清华大学; 中国科学院自动化研究所; 中关村学院; 南开大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TrustMOPD,一种无标签的令牌级多教师在线策略蒸馏方法,通过可靠性校准加权监督,显著提升恢复比率并接近标签基线的性能。
AI 中文摘要
多教师在线策略蒸馏允许学生模型在其自身的轨迹上向互补的专家学习。然而,基于领域路由的方法为每个示例选择一个教师,并在整个响应过程中保持该教师不变。这种设计既依赖于混合训练语料库通常缺乏的标签,又无法在轨迹内所需专业知识发生变化时调整教师选择。我们提出TrustMOPD,用无标签的、令牌级别的监督分配取代示例级别的教师选择。在每个学生生成的前缀处,TrustMOPD使用每个专家从共享的预强化学习参考中产生的强化学习引起的位移作为局部可靠性的代理,跨教师校准这些分数,并构建加权蒸馏目标。在数学、代码和指令跟随任务中,TrustMOPD优于最强的无标签基线,将SingleCap上的恢复比率从54.4%提高到91.5%,将MultiCap上的恢复比率从54.5%提高到98.0%,同时在SingleCap上接近基于标签的MOPD。独立于学生生成的前缀随机化令牌级权重并不比均匀加权更好,这支持了将监督条件化于不断演化的生成上下文的重要性。
英文摘要
Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, however, select one teacher per example and keep it fixed throughout the response. This design both depends on domain labels that mixed training corpora often lack and cannot adapt teacher selection when the expertise required changes within a trajectory. We observe that each specialist deviates more from a shared reference on in-domain prompts than on out-of-domain prompts, on average. Based on this observation, we propose \textbf{TrustMOPD}, which replaces example-level teacher selection with label-free, token-level supervision allocation. At each student-generated prefix, TrustMOPD measures this displacement in next-token preferences, calibrates its magnitude across teachers, and uses the resulting scores as proxies for local reliability to weight teacher-specific distillation losses. Evaluated across mathematics, code, and instruction following, TrustMOPD closes 91.5\% and 98.0\% of the overall-score gap between the initial student and oracle-routed teachers when trained on \textsc{SingleCap} and \textsc{MultiCap}, respectively, compared with 54.4\% and 54.5\% for the strongest label-free baseline in each setting. On \textsc{SingleCap}, it approaches label-based MOPD without using domain labels.
Comments33 pages, 9 figures, 12 table