arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Video-MOPD:面向视频理解的多教师在线策略蒸馏

Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding

Zhenxin Qin, Peng Shi, Cong Han, Yinlong Qian, Zequn Jie, Lin Ma

arXiv 2609.09300首次发表:更新:

发表机构

Tongji University; Bilibili Inc.(同济大学; 哔哩哔哩股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Video-MOPD-8B,通过多教师在线策略蒸馏和可靠性感知采样,在视频时间定位、通用理解和STEM推理上实现全面最先进性能。

AI 中文摘要

视频理解要求感知、时间理解和复杂推理等互补能力的融合,这些能力难以在单一模型中联合优化。我们推出了Video-MOPD-8B,一个专用于视频理解任务的开源权重模型。为从根本上增强其能力,我们在三个核心领域进行了针对性的强化学习(RL)优化:视频时间定位(VTG)、通用视频理解和视频STEM推理。随后,我们通过多教师在线策略蒸馏(MOPD)统一这些互补能力,该方法通过路由教师反馈监督学生生成的轨迹来整合专家知识。我们进一步引入了可靠性感知信息采样(RAIS),该采样方法选择具有一致可靠教师监督且师生性能差距较大的样本。这些组件共同使Video-MOPD-8B能够在多样化的视频理解任务中实现协调且全面的性能提升。在涵盖通用视频理解、时间定位、视频推理和视频STEM任务的综合基准上的大量实验表明,Video-MOPD-8B在同等规模现有模型中达到了最先进的性能。训练好的模型权重可在该https URL获取。

英文摘要

Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are difficult to jointly optimize within a single model. We introduce Video-MOPD-8B, an open-weight model dedicated to video understanding tasks. To fundamentally enhance its capabilities, we conduct targeted reinforcement learning (RL) optimization across three core domains: video temporal grounding (VTG), general video comprehension, and video STEM reasoning. We then unify their complementary capabilities via Multi-Teacher On-Policy Distillation (MOPD), which consolidates expert knowledge by supervising student-generated trajectories with routed teacher feedback. We further introduce Reliability-Aware Informative Sampling (RAIS), which selects examples with consistently reliable teacher supervision and large teacher-student performance gaps. Together, these components enable Video-MOPD-8B to achieve coordinated and comprehensive performance gains across diverse video understanding tasks. Extensive experiments on comprehensive benchmarks covering general video understanding, temporal grounding, video reasoning, and video STEM tasks demonstrate that Video-MOPD-8B achieves state-of-the-art performance among existing models at a comparable scale. The trained model weights are available at https://huggingface.co/LandH/Video-MOPD-8B.

CommentsTechnical report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑