arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VisualRouter:面向长视频理解的查询驱动视觉采样

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding

Haiyue Zhang, Yi Bin, Xun Jiang, Zeyu Ma, Duo Peng, Guoqing Wang, Yang Yang, Heng Tao Shen

arXiv 2607.28463首次发表:更新:

发表机构

Tongji University; University of Electronic Science and Technology of China(同济大学; 电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出无训练即插即用框架VisualRouter,通过将查询分为全局或局部并采用对应采样策略,提升了大型视觉语言模型的长视频理解性能,在多个基准数据集上优于均匀采样及现有无训练方法。

AI 中文摘要

大型视觉语言模型(LVLMs)在视频理解领域已取得显著进展,但由于视觉token数量庞大且上下文窗口有限,长视频理解仍面临挑战。视觉采样通过选择信息丰富的帧子集提供了实用解决方案,不过现有方法通常要么依赖关联感知采样,导致帧选择冗余且时间覆盖不足,要么采用与查询类型无关的固定采样策略。本文提出VisualRouter,这是一种无训练、即插即用的查询驱动视觉采样框架。VisualRouter首先将每个查询分类为全局或局部,再应用对应采样策略:针对全局查询,采用关联-覆盖混合策略,在保留时间覆盖的同时保留与查询相关的视觉证据;针对局部查询,采用事件感知帧选择策略,执行事件划分、段级帧分配和事件内帧选择,在输入帧数量有限的情况下共同平衡关联度、覆盖度和多样性。实验表明,VisualRouter在多个大型视觉语言模型上相比均匀采样均有提升,使用Qwen2.5-VL-7B时在Video-MME、LongVideoBench和MLVU上分别实现5.2%、7.7%和11.6%的提升,且在相同设置下优于现有无训练视觉采样方法。

英文摘要

Large vision-language models (LVLMs) have achieved significant progress in video understanding, yet understanding long videos remains challenging due to the large number of visual tokens and limited context windows. Visual sampling provides a practical solution by selecting an informative subset of frames. However, existing methods typically either rely on relevance-aware sampling, leading to redundant frame selection and insufficient temporal coverage, or adopt a fixed sampling strategy regardless of query type. In this paper, we propose VisualRouter, a training-free and plug-and-play framework for query-grounded visual sampling. VisualRouter first classifies each query as either global or local and then applies the corresponding sampling strategy. For global queries, it employs a relevance-coverage hybrid strategy that preserves temporal coverage while retaining query-relevant visual evidence. For local queries, it adopts an event-aware frame selection strategy that performs event partitioning, segment-level frame allocation, and intra-event frame selection, jointly balancing relevance, coverage, and diversity with a limited number of input frames. Experiments show that VisualRouter consistently improves multiple LVLMs over uniform sampling, achieving gains of 5.2%, 7.7%, and 11.6% on Video-MME, LongVideoBench, and MLVU with Qwen2.5-VL-7B, and outperforming existing training-free visual sampling methods under the same setting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑