arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越教师分配:域归一化多教师同策略蒸馏

Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Xin Li, Hao Jiang, Xin Gao, Annan Wang, Yuchen Xie, Jinghao Guo, Xingwei Qu, Yichi Zhang, Chau Yuen

arXiv 2609.35347首次发表:更新:

AI 中文总结

针对多教师蒸馏中反馈不平衡问题,提出域归一化MOPD,按领域反馈分散度重缩放,在六个基准上提升平均分并恢复数学增益。

AI 中文摘要

强化学习可以将一个语言模型转化为多个专家,每个专家擅长单一技能,如数学、编码或指令遵循,但用户需要的是一个具备所有这些技能的单一模型。多教师同策略蒸馏(MOPD)通过让专家教导一个学生模型来合并这些技能:学生回答每个提示,而该提示所属领域的专家对每个词元提供反馈。这种路由机制决定了哪位专家进行教学,但并未决定其反馈对共享学生模型的更新强度有多大。在三个规模的Qwen3.5模型中,我们发现MOPD的学生模型并未超越由最佳单一专家教导的学生模型,并且几乎未能获得数学专家的优势。反馈是不平衡的:指令遵循反馈的分散程度是数学反馈的数倍,并主导了学生模型的更新。我们提出了域归一化MOPD(DN-MOPD),该方法保留路由机制,并根据各领域反馈的实测分散程度对其重新缩放。在六个公开基准上,DN-MOPD在每种规模、三种随机种子和两种答案长度限制下,均比MOPD提高了平均得分,并恢复了大部分丢失的数学增益。使用固定域权重的对照实验表明,增益主要来自降低指令遵循反馈的权重,而非单独提高数学反馈的权重,且接近DN-MOPD所测得的固定权重表现相当。因此,组合专家不仅需要决定由哪位专家教学,还需要决定其反馈的权重有多大。

英文摘要

Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.

CommentsProject page: https://lixin.ai/DN-MOPD . Code: https://github.com/LiXin97/DN-MOPD

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑