发表机构
Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在提升全模态模型能力,提出基于策略的全模态蒸馏(OPOD)方法,将学生响应路由到对应模态教师,调整教师影响,训练后仅留可部署的全模态模型,在多基准测试和多种骨干网络规模下取得优异成绩。
AI 中文摘要
全模态模型能在一个系统中处理文本、图像和音频,但同时提升所有这些能力仍很困难。在聚合多模态数据上训练单个模型往往无法与针对单个模态的专门模型相匹配。基于策略的蒸馏(OPD)提供了一种整合此类专门模型的方法,但使用多个教师会引入相互竞争的指导并以牺牲一种模态为代价来提升另一种模态。我们提出基于策略的全模态蒸馏(OPOD),它将每个学生响应路由到匹配的文本、图像或音频教师。OPOD仅在教师分配概率高于学生的令牌上保留教师指导,在训练期间独立调整每个模态教师的影响,并要求路由后的教师评估最终答案以及推理是否支持正确答案。在十二个基准测试和三种骨干网络规模上,OPOD在每个规模上都取得了最佳平均分数,在30B模型上,在所有十二个基准测试中均优于基础模型和在聚合多模态数据上联合后训练的对应模型,甚至在包含单个专门模型时,在十一个基准测试中排名第一或第二。训练后丢弃专门模型,留下一个可部署的全模态模型。这些结果表明,协调特定模态的教师是在保持跨模态平衡的同时改进共享模型的有效方法。
英文摘要
Omni-modal models provide a unified interface for text, images, and audio. However, improving these abilities together remains difficult, as post-training on pooled multimodal data often fails to preserve the strengths of modality teachers. On-policy distillation (OPD) has recently become popular in model post-training. It samples responses from the current student and compares the teacher's and student's next-token distributions along those responses, yielding dense supervision while reducing the mismatch between training and inference. Despite these advantages, standard OPD does not readily extend to several modality teachers. Their guidance may favor conflicting changes to the shared model, while matching each teacher's next-token distribution can prevent the student from moving beyond that teacher. To address these challenges, we propose On-Policy Omni Distillation (OPOD), which consolidates text, image, and audio teachers into one omni model. OPOD routes each response to the corresponding teacher, controls the teachers independently, and applies guidance only when the teacher assigns a higher probability to the generated token. The selected teacher also evaluates answer confidence and whether the reasoning increases support for the answer. Extensive experiments on twelve benchmarks show that OPOD achieves the best average at three model scales, reaching 70.8, 51.7, and 46.2 and outperforming the strongest comparator by 2.1, 1.8, and 1.7 points. At 30B, it surpasses the base model and pooled RL training on all twelve benchmarks, and ranks first or second on eleven even when the teachers are included. Only the student is retained for deployment.