arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Video-OPSD:利用特权视觉证据实现视频大语言模型中的在线自蒸馏

Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models

Ziyue Wang, Shiqi Huang, Weiwen Xu, Bihan Wen, Xudong Jiang

arXiv 2608.27065首次发表:更新:

发表机构

School of Electrical and Electronic Engineering, Nanyang Technological University; The Chinese University of Hong Kong(南洋理工大学电气与电子工程学院; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Video-OPSD框架,利用视频内特权视觉证据构建自教师并优化令牌级蒸馏,在多个视频基准上优于标准OPSD,性能接近GRPO且训练时间大幅缩短,为Video-LLMs提供高效后训练方法。

AI 中文摘要

在线自蒸馏(OPSD)是一种新兴的有效后训练范式,通过特权自教师提供的密集令牌级监督来改进策略优化。尽管OPSD前景良好,但它在视频大语言模型(Video-LLMs)中的研究仍十分有限。现有方法通常通过为教师增加额外信息来构建特权教师,同时保持教师和学生的主要输入不变。然而,视频推理可在主要输入内部提供独特的特权监督源:长视频包含大量时间冗余,仅一小部分帧提供回答问题所需的证据。基于此观察,我们提出Video-OPSD,这是一种利用特权视觉证据进行自教师构建和知识迁移的OPSD框架。首先,我们的证据基础自教师仅基于带注释的证据帧条件化教师,而学生继续对完整视频进行推理,这种聚焦的视觉输入使教师能提供更具信息性的监督。其次,我们的证据引导令牌优化根据每个推理令牌对特权视觉证据的依赖程度自适应加权令牌级蒸馏,从而强调基于感知的推理。在多个视频理解和推理基准上的实验表明,Video-OPSD在多个主干网络上始终优于标准OPSD,且达到与GRPO相当的性能,同时所需训练时间显著更少,为Video-LLMs建立了一种有效且高效的后训练方法。

英文摘要

On-policy self-distillation (OPSD) has recently emerged as an effective post-training paradigm that improves policy optimization through dense token-level supervision from a privileged self-teacher. Despite its promise, OPSD remains largely underexplored for Video Large Language Models (Video-LLMs). Existing methods typically construct privileged teachers by augmenting their context with additional information while keeping the primary input unchanged for both teacher and student. Video reasoning, however, offers a distinct source of privileged supervision within the primary input itself: long videos contain substantial temporal redundancy, and only a small subset of frames provides the evidence necessary to answer a question. Building on this observation, we present $\textbf{Video-OPSD}$, an OPSD framework that exploits privileged visual evidence for both self-teacher construction and knowledge transfer. First, our Evidence-Grounded Self-Teacher conditions the teacher exclusively on annotated evidence frames while the student continues to reason over the complete video. This focused visual input enables the teacher to provide more informative supervision. Second, our Evidence-Guided Token Optimization adaptively weights token-level distillation according to each reasoning token's reliance on privileged visual evidence, thereby emphasizing perceptually grounded reasoning. Experiments across video understanding and reasoning benchmarks show that $\textbf{Video-OPSD}$ consistently improves upon Standard OPSD across multiple backbones and achieves performance comparable to GRPO while requiring substantially less training time, establishing an effective and efficient post-training approach for Video-LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑