arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

向任何正确的对象学习:用于多领域大语言模型的答案验证多教师蒸馏

Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs

Xixiang He, Xingming Li, Baiqi Wu, Qiyao Sun, Xuanyu Ji, Ao Cheng, Qingyong Hu

arXiv 2609.02548首次发表:更新:

发表机构

National University of Defense Technology; Zhejiang University; Intelligent Game and Decision Lab(国防科技大学; 浙江大学; 智能游戏与决策实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MT-SDPO方法,通过答案验证识别可靠教师而非依赖领域,将多教师蒸馏统一为单模型,提升了Qwen3-8B的领域性能并缩小了差距。

AI 中文摘要

现代大语言模型(LLMs)依赖强化学习构建各领域的强大能力,但将这些能力整合为单个可部署模型仍具挑战。现有方法通过将每个样本路由至领域匹配的教师,让领域标签决定哪个教师提供监督。然而,领域专业知识仅在平均层面成立:匹配的教师在给定样本上并不总是正确,而来自另一领域的教师有时是正确的。因此,必须针对每个样本而非每个领域识别可靠教师。本文提出多教师自蒸馏策略优化(MT-SDPO),这是一种在线蒸馏方法,可将多个冻结教师统一为一个学生模型。MT-SDPO包含三个组件:(1)自锚点,即一次rollout由其自身组中的正确rollout进行监督;(2)答案验证资格,即教师仅在自身答案通过验证器时才可监督样本;(3)特权蒸馏,即将锚点和所有验证反馈合并为一个上下文,指数移动平均自教师读取该上下文,而学生不读取,从而在部署时保持一个策略。在来自三个模型家族的五个学生模型上,MT-SDPO将Qwen3-8B的最弱领域提升了14.79个百分点,并将其领域差距缩小了74.7%,比每个领域服务一个匹配教师的方法实现了更好的平衡。决定谁来教学的应是经过验证的可靠性,而非领域成员身份。代码可在该https URL获取。

英文摘要

Modern large language models (LLMs) rely on reinforcement learning to build strong capabilities in individual domains, but integrating those capabilities into a single deployable model remains challenging. By routing each sample to the teacher whose domain matches it, existing approaches let a domain label decide which teacher provides supervision. However, domain expertise holds only on average: the matched teacher is not always correct on a given sample, while a teacher from another domain sometimes is. The reliable teacher therefore has to be identified per sample, not per domain. In this paper, we introduce Multi-Teacher Self-Distillation Policy Optimization (MT-SDPO), an on-policy distillation method that unifies several frozen teachers into one student model. MT-SDPO consists of three components: (1) self-anchors, where a rollout is supervised by a correct rollout from its own group; (2) answer-verified eligibility, where a teacher may supervise a sample only if its own answer passes a verifier; and (3) privileged distillation, which merges the anchor and all verified feedback into one context that an exponential moving average self-teacher reads and the student does not, thereby keeping one policy at deployment. Across five students from three model families, MT-SDPO lifts the weakest domain of Qwen3-8B by 14.79 points and narrows its domain gap by 74.7%, a better balance than serving one matched teacher per domain. Verified reliability, not domain membership, should decide who teaches. Code is available at https://github.com/hexixiang/MT-SDPO.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑