arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MOPD-Router:在多教师在线策略蒸馏中重新思考教师路由

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma, Jihuai Zhu, Lelun Wu, Lyumanshan Ye, Pengfei Liu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu

arXiv 2609.30837首次发表:更新:

发表机构

Shanghai Jiao Tong University; GAIR; Alibaba Group; University of Science and Technology of China; Shanghai Innovation Institute(上海交通大学; GAIR; 阿里巴巴集团; 中国科学技术大学; 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MOPD-Router框架,通过token级路由和ExpertAlign指标,无需领域标签即可在多教师在线策略蒸馏中有效利用跨领域互补监督,提升学生模型性能。

AI 中文摘要

多教师在线策略蒸馏(MOPD)将多个教师的专长整合到单个学生模型中,但现有实践通常将每个提示硬路由到领域匹配的教师进行整个轨迹生成。这种对提示级领域标签的依赖限制了未标注训练混合数据的使用,并使得其他教师的互补信号未被利用。我们提出MOPD-Router,一个在每个token上从整个教师池中路由监督信号的框架,无需领域标签或训练单独的路由模型。其即插即用接口支持不同的指标来选择和加权教师特定的在线策略蒸馏(OPD)信号。在该接口内,我们提出ExpertAlign,它根据教师对当前token的学生修正是否表达了该教师在后期训练中获得的专长来为每个教师打分,并将其与基于教师置信度(熵)和教师-学生差异(新颖性)的两个参考指标进行比较。在强到弱和同规模蒸馏场景下的未标注和领域标注训练混合数据上的实验表明,ExpertAlign在所有四种设置中均取得了最强的整体性能。在未标注数据上,它比均值聚合的整体得分提高了5.88(+12.3%)分;在领域标注数据上,它比标准MOPD提高了3.95(+7.8%)分,且未使用可用的领域标签。这些结果表明,token级路由能够利用跨领域的互补监督,并减少对提示级领域分配的依赖。代码可在以下网址获取:此https URL。

英文摘要

Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model. Its plug-in interface supports different metrics for selecting and weighting teacher-specific OPD signals. Within this interface, we propose ExpertAlign, which scores each teacher by whether its correction to the student at the current token expresses the specialization that teacher acquired during post-training, and compare it against two reference metrics built on teacher confidence (Entropy) and teacher-student discrepancy (Novelty). Experiments on unlabeled and domain-labeled training mixtures under strong-to-weak and same-size distillation scenarios show that ExpertAlign achieves the strongest overall performance in all four settings. On unlabeled data, it improves the overall score by 5.88 (+12.3%) points over Mean aggregation; on domain-labeled data, it outperforms standard MOPD by 3.95 (+7.8%) points without using available domain labels. These results demonstrate token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment. Code is available at: https://github.com/TURLEing/MOPD-Router.

Comments19 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑