arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于时间视频定位的反事实注意力策略蒸馏

Counterfactual Attention Policy Distillation for Temporal Video Grounding

Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Jun Peng, Yiyi Zhou

arXiv 2609.34581首次发表:更新:

发表机构

Xiamen University; Fudan University; The Chinese University of Hong Kong; University of the Chinese Academy of Sciences(厦门大学; 复旦大学; 香港中文大学; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长视频中重复动作和相似上下文导致的时间定位困难,提出反事实注意力策略蒸馏(CAPD),通过掩蔽时间组校准教师注意力并加权蒸馏,在Qwen3-VL-8B上以2500样本训练,使TimeLens平均召回率相对提升12.0%。

AI 中文摘要

时间视频定位是先进多模态大语言模型(MLLMs)理解视频事件的关键能力,然而在长视频中,重复动作和视觉相似上下文常常限制其性能。本文从在线策略蒸馏(OPD)的角度研究该问题,并提出一种新的MLLMs训练机制,称为反事实注意力策略蒸馏(CAPD)。具体而言,OPD通过为学生生成的轨迹提供密集的教师监督,是MLLMs的可行解决方案。但其基于下一令牌的师生蒸馏难以识别支持每个预测时间戳的特定视频片段,而这对于时间定位至关重要。在这种情况下,CAPD通过衡量掩蔽每个时间组如何改变教师的输出分布来工作。由此产生的反事实影响校准教师的注意力并加权令牌级蒸馏,使学生能够学习影响边界预测的时间证据。为验证CAPD,我们在Qwen3-VL-8B-Instruct上仅使用2500个样本训练一个周期,并在TimeLens和多个通用视频基准上进行评估。实验结果表明,与GRPO相比,CAPD在TimeLens上将平均召回率相对提升了12.0%,同时保持通用视频理解能力,达到与基础模型相当的准确率。

英文摘要

Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of On-policy distillation (OPD) and propose a new training regime for MLLMs termed Counterfactual Attention Policy Distillation (CAPD). In particular, OPD is a viable solution for MLLMs via providing dense teacher supervision on student-generated trajectories. But its next-token based teacher-student distillation is hard to identify the specific video segments supporting each predicted timestamp, which is critical for temporal grounding. In this case, CAPD measures how masking each temporal group changes the teacher's output distribution. The resulting counterfactual influence calibrates the teacher's attention and weights token-level distillation, allowing the student to learn the temporal evidence that affects boundary prediction. To validate CAPD, we trained it on Qwen3-VL-8B-Instruct using only 2,500 samples for one epoch, and evaluated it on the TimeLens and multiple general video benchmarks. Experimental results show that CAPD improves average recall by 12.0% relative to GRPO on TimeLens while preserving general video understanding, achieving comparable accuracy to the base model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑