arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15778cs.CVcs.AI

用于多事件长视频理解的模块化动态粒度视频语言模型

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

Wei Feng, Xin Wang, Yu-Wei Zhan, Yuwei Zhou, Wenwu Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

针对长视频理解中视觉令牌预算与多事件捕捉的矛盾问题,提出MoD-VLLM框架,含正-负视频片段定位和模块化动态粒度反射模块,结合动态粒度强化学习策略,在多事件长视频基准测试中显著优于现有基线。

中文摘要 AI 辅助

视频大语言模型(Video LLMs)在各种视频理解任务中取得了显著进展。然而,由于视觉令牌预算有限与捕获多个关键事件的需求之间的矛盾,长视频场景仍然具有挑战性。现有方法通常分两个阶段处理长视频,存在缺乏自适应能力分配和自我校正的模块化机制等局限性。为应对这些挑战,我们提出了MoD-VLLM,一种用于多事件长视频理解的新颖的模块化动态粒度视频语言模型框架,它迭代地和自我反思地统一了时间定位和语义理解。具体而言,我们提出了正-负视频片段定位模块和模块化动态粒度反射模块,它们形成一个闭环以逐步定位与问题相关的视频片段。定位模块根据视频问题指导视频语言模型区分相关和不相关的视频片段。反射模块采用模块化调度器,为相关正片段动态选择细粒度编码以捕获详细感知,为负片段选择粗粒度编码以维持全局上下文。我们还提出了一种动态粒度强化学习策略,使MoD-VLLM能够联合学习最优定位策略和动态粒度视觉表示。此外,我们提出了MEventBench,一个用于复杂长视频推理的具有挑战性的多事件长视频基准。在几个长视频理解基准和我们的MEventBench上进行的大量实验表明,MoD-VLLM显著优于现有最先进的基线。

英文摘要

Video Large Language Models (Video LLMs) have made significant advancements in various video understanding tasks. However, long-video scenarios remain challenging due to the tension between limited visual token budgets and the need to capture multiple key events. Existing approaches typically process long videos in two stages, i.e., i) select keyframes and ii) perform detailed perception, which exhibit limitations: they lack a modular mechanism for adaptive capacity allocation and self-correction, resulting in unreliable modeling. To tackle these challenges, we propose MoD-VLLM, a novel Modularized Dynamic-Granularity Video LLM framework for multi-event long video understanding, which unifies temporal grounding and semantic understanding iteratively and self-reflectively. Specifically, we propose a Positive-Negative Video Segments Grounding module and a Modularized Dynamic-Granularity Reflection module, which form a closed loop to progressively localize the question-related video segments. The grounding module instructs a Video LLM to distinguish relevant from irrelevant video segments based on the video question. The reflection module employs a modularized scheduler that dynamically selects fine-grained encoding for relevant positive segments to capture detailed perception and coarse-grained encoding for negative segments to maintain global context. We further propose a dynamic-granularity reinforcement learning strategy, allowing MoD-VLLM to learn optimal grounding policies and dynamic granularity visual representation jointly. Moreover, we propose MEventBench, a challenging Multi-Event Long Video Benchmark for complex long video reasoning. Extensive experiments on several long video understanding benchmarks and our MEventBench demonstrate that MoD-VLLM significantly outperforms state-of-the-art baselines.

补充信息

↑