arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

H-OPD:置信感知异构多教师多模态在线策略蒸馏

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation

Qixiang Yin, Huanjin Yao, Yuchen Cai, Jianghao Chen, Ziyi Wang, Min Yang, Fei Su, Zhicheng Zhao

arXiv 2607.02592首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; ByteDance; USTC; Beijing Key Laboratory of Network System and Network Culture; Key Laboratory of Interactive Technology and Experience System, Ministry of Culture and Tourism; Zhongguancun Academy(北京邮电大学; 字节跳动; 中国科学技术大学; 北京网络系统与网络文化重点实验室; 文化和旅游部互动技术与体验系统重点实验室; 中关村科学城)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多模态推理的在线策略蒸馏问题,提出H-OPD框架,通过验证异构教师互补性,以令牌级教师仲裁取代任务或样本级路由,结合视觉与文本教师,在多基准测试中性能优越。

AI 中文摘要

在线策略蒸馏(OPD)是一种有效的训练后范式。现有多模态推理的OPD方法依赖静态教师路由,忽略不同解码步骤主导模态不同。为此提出H-OPD,通过验证异构教师互补性,用令牌级教师仲裁取代任务或样本级路由,采用视觉到语言描述转移,在各令牌处动态组合教师,多基准测试验证了其优越性能。

英文摘要

On-policy distillation (OPD) has recently emerged as an effective post-training paradigm by providing supervision on student-generated trajectories. However, existing OPD methods for multimodal reasoning usually rely on a static teacher routing, assigning each sample to a single teacher based on modality or task type. This ignores that visual grounding and abstract reasoning may dominate different decoding steps, making a single teacher insufficient for the full trajectory. To this end, H-OPD is proposed as a confidence-aware heterogeneous multi-teacher OPD framework for multimodal reasoning. By verifying the complementarity of heterogeneous teachers in the same reasoning process, H-OPD replaces task or sample level teacher routing with token-level teacher arbitration along the shared student trajectory. H-OPD employs vision-to-language description transfer to enable text-only teachers to access key visual semantics, and uses a confidence-aware arbitration mechanism to dynamically combine vision-language teacher and text-only teachers at each token. Extensive evaluations over 11 widely-used reasoning benchmarks showcase the superior performance of our method.

CommentsEMNLP2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑