arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.02994cs.CV

Video-OPD:通过在线策略蒸馏实现多模态大语言模型在时序视频定位中的高效后训练

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

Jiaze Li, Hao Yin, Haoran Xu, Boshen Xu, Wenhui Tan, Zewen He, Jianzhong Ju, Zhenbo Luo, Jian Luan

首次发表 更新
浏览论文内容

中文总结 AI 辅助

提出Video-OPD框架,利用在线策略蒸馏和教师验证分歧聚焦课程,以高效后训练多模态大语言模型进行时序视频定位,克服稀疏奖励和高计算开销问题。

中文摘要 AI 辅助

强化学习因其在线策略优化而成为时序视频定位(TVG)后训练的一种有原则的范式,但现有的基于GRPO的方法仍然受到稀疏奖励信号和大量计算开销的根本限制。我们提出了Video-OPD,一个受近期在线策略蒸馏进展启发的TVG高效后训练框架。Video-OPD优化直接从当前策略采样的轨迹,从而保持训练和推理分布之间的一致性,同时前沿教师通过反向KL散度目标提供密集的令牌级监督。这种公式保留了缓解分布偏移至关重要的在线策略属性,同时将稀疏的回合级反馈转化为细粒度的逐步学习信号。基于Video-OPD,我们引入了教师验证分歧聚焦(TVDF),一种轻量级训练课程,迭代地优先考虑既教师可靠又对学生信息量最大的轨迹,从而提高训练效率。实验结果表明,Video-OPD在实现显著更快的收敛和更低计算成本的同时,始终优于GRPO,确立了在线策略蒸馏作为TVG传统强化学习的有效替代方案。

英文摘要

Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet existing GRPO-based methods remain fundamentally constrained by sparse reward signals and substantial computational overhead. We propose Video-OPD, an efficient post-training framework for TVG inspired by recent advances in on-policy distillation. Video-OPD optimizes trajectories sampled directly from the current policy, thereby preserving alignment between training and inference distributions, while a frontier teacher supplies dense, token-level supervision via a reverse KL divergence objective. This formulation preserves the on-policy property critical for mitigating distributional shift, while converting sparse, episode-level feedback into fine-grained, step-wise learning signals. Building on Video-OPD, we introduce Teacher-Validated Disagreement Focusing (TVDF), a lightweight training curriculum that iteratively prioritizes trajectories that are both teacher-reliable and maximally informative for the student, thereby improving training efficiency. Empirical results demonstrate that Video-OPD consistently outperforms GRPO while achieving substantially faster convergence and lower computational cost, establishing on-policy distillation as an effective alternative to conventional reinforcement learning for TVG.

发表机构

  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑