发表机构
Ant Security Lab, Ant Group; Fudan University(蚂蚁安全实验室,蚂蚁集团; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多教师蒸馏导致学生模型通用能力下降的问题,提出SF-MOPD方法,通过快速与慢速模型耦合,仅移除偏离慢速模型的更新成分,有效缓解能力干扰并提升多模态专长。
AI 中文摘要
基础多模态大语言模型旨在支持跨多个领域的广泛能力。多教师在线蒸馏(MOPD)提供了一种有效的框架,将特定领域的专业知识整合到单个学生模型中。然而,MOPD训练逐渐使学生模型偏离其初始化模型,随着偏移量的增加,通用能力下降,导致能力干扰。直接的补救措施是约束学生模型向其初始化靠拢,但这同样抑制了领域专业知识的获取。我们提出了慢-快多教师在线蒸馏(SF-MOPD),该方法将一个快速模型(即由每个教师直接更新的当前学生模型)与一个慢速模型(即学生模型的指数移动平均)相结合。慢速模型逐渐吸收学习信号,作为一个移动的能力参考,融合了通用基础与已确认的领域专业知识。对于每个教师,SF-MOPD在对数概率空间中计算教师诱导的更新,并仅移除将快速模型推离慢速模型更远的成分,同时保留对齐和正交的成分。跨多个模型规模的实验表明,SF-MOPD有效缓解了能力干扰,增强了专业多模态能力,并减少了通用能力基准上的平均退化,始终优于原始的MOPD。
英文摘要
Foundation multimodal large language models are designed to support a broad spectrum of capabilities across diverse domains. Multi-teacher on-policy distillation (MOPD) provides an effective framework for consolidating domain-specific expertise into a single student model. However, MOPD training gradually drives the student away from its initialization model, and general capabilities decline as the displacement grows, resulting in capability interference. A direct remedy is constraining the student toward its initialization, but this suppresses the acquisition of domain expertise as well. We propose Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model, the current student updated directly by each teacher, with a slow model, an exponential moving average of the student. The slow model absorbs the learning signal gradually, serving as a moving capability reference that fuses the general foundation with confirmed domain expertise. For each teacher, SF-MOPD computes the teacher-induced update in log-probability space and removes only the component that pushes the fast model further away from the slow model, while retaining aligned and orthogonal components. Experiments across multiple model scales demonstrate that SF-MOPD effectively mitigates capability interference, enhances specialized multimodal capabilities, and reduces the average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.
Comments5 pages, 2 figures