arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05592cs.CV

超越帧选择:用多模态大语言模型重新思考长视频理解

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

Ziling Huang, Yuki M. Asano, Shin'ichi Satoh

首次发表
浏览论文内容

中文总结 AI 辅助

针对 MLLMs 长视频理解的帧选择策略难以兼顾全局与局部信息的问题,提出 VideoRouter 模型,通过时间层次结构和验证引导路由器协调全局与局部推理,在 VideoMME 数据集上实现性能提升。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)在视频理解领域已取得显著进展,但受限于 token 数量,MLLMs 难以捕捉时间上稀疏的证据,这仍是一大挑战。现有方法通常依赖均匀采样或帧选择策略,这些策略往往要么优化广泛的时间覆盖,要么优化局部相关性,难以同时保留全局剧情上下文和细粒度证据。我们提出 VideoRouter(VR),将长视频理解重新定义为协调互补的证据视图,而非选择单一帧子集。该方法首先将每个视频组织为与问题无关的时间层次结构,将视频划分为从粗到细的时间连贯片段;在该层次结构中,上层节点捕获广泛的剧情上下文和事件进展,下层节点保留细粒度的局部细节和证据时刻,由此自然产生两个互补视图:面向覆盖范围推理的全局视图,以及面向细节证据恢复的局部视图。我们进一步引入验证引导路由器,以确定所选证据更支持哪个视图并选择最终答案。我们通过大量实验验证了该设计的有效性,结果显示,在 LLaVA-Video-7B 骨干网络下,所提方法在 VideoMME 数据集上较最先进的帧选择方法提升了 2.9 个百分点,我们将发布代码。

英文摘要

Multimodal Large Language Models (MLLMs) have made strong progress in video understanding, yet long videos remain difficult: the visual token budget grows with video length, so temporally sparse evidence is easily lost. Existing methods, whether frame selection or agent-based exploration, ultimately rely on a single frame set that must trade off broad temporal coverage against fine-grained, question-relevant evidence: mixing both dilutes decisive evidence, while focusing on either sacrifices the other. We propose VideoRouter (VR), which rethinks long-video understanding as coordinating complementary evidence views rather than selecting a single subset of frames. VideoRouter first organizes each video into a question-agnostic temporal hierarchy that partitions it into coarse-to-fine temporally coherent segments. Upper-level nodes capture broad storyline context and event progression, while lower-level nodes preserve fine-grained local details and evidence-bearing moments. This gives rise to two complementary views: a global view for coverage-oriented reasoning and a local view for detail-oriented evidence recovery. We further introduce a verification-guided router that judges which view is better supported by its own selected evidence and decides the final answer. Across six backbones, routing improves over both views in all settings, and the choice of view is shown to be dataset-dependent, confirming that no single evidence granularity is universally preferable. On VideoMME and LongVideoBench, our method outperforms state-of-the-art frame-selection methods by 2.5 and 1.3 points, respectively, under the LLaVA-Video-7B backbone. We will release the code.

发表机构

  • National Institute of Informatics(信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑