发表机构
University of Science and Technology of China; USTC; Suzhou Institute for Advanced Research(中国科学技术大学; 中国科学技术大学; 苏州高等研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VideoMM提出自适应宏-微推理范式,通过解耦选择与推理,在宏代理上过滤语义区域并仅在必要时使用微标记,实现长视频理解中6.13倍加速与7.4%准确率提升。
AI 中文摘要
将多模态大语言模型(MLLMs)扩展到长视频理解受到视觉标记爆炸的瓶颈限制,这会使上下文窗口饱和并产生高昂成本。现有解决方案主要依赖辅助模型进行标记缩减,但面临根本性困境:轻量级编码器驱动的方法常常忽略关键语义信息,而重量级MLLM驱动的缩减则抵消了效率提升。在这项工作中,我们识别出这一困境背后更根本的低效问题:虽然细粒度的视觉细节对于详细理解至关重要,但在初步选择语义相关区域的任务中,这些细节在很大程度上是冗余的。受此启发,我们提出了VideoMM,它标志着从以模型为中心的缩减向自适应感知粒度的范式转变。具体而言,我们的框架将选择与推理解耦,通过在成本效益高的宏代理(由降采样帧派生)上执行语义过滤,并仅在必要时将选定区域投影到高保真微标记上进行详细理解。大量评估表明,VideoMM显著优于现有解决方案。在LongVideoBench上,它比全上下文基线实现了6.13倍的加速和7.4%的准确率提升,并进一步将推理速度比当前领先方法提升2.73倍,为长视频理解建立了一种高度可扩展的范式。我们的代码可在以下网址获取:this https URL。
英文摘要
Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight encoder-driven approaches often overlook critical semantic information, whereas heavyweight MLLM-driven reduction negates the efficiency gains. {In this work, we identify a more fundamental inefficiency underlying this dilemma: while fine-grained visual details are essential for detailed understanding, they are largely redundant for the preliminary task of selecting semantically relevant regions. } Motivated by this, we introduce \textbf{VideoMM}, which marks a paradigm shift from model-centric downsizing to adaptive perceptual granularity. Specifically, our framework {decouples selection from reasoning} by executing semantic filtering on a cost-effective \textit{Macro Proxy} (derived from downscaled frames), and projecting the selected regions onto high-fidelity \textit{Micro Tokens} for detailed understanding only when necessary. Extensive evaluations show that VideoMM significantly outperforms existing solutions. It achieves a 6.13$\times$ speedup and a 7.4\% accuracy gain over full-context baselines on LongVideoBench, and further accelerates inference by 2.73$\times$ over current leading methods, establishing a highly scalable paradigm for long-video understanding. Our code is available at: https://github.com/adfh917k/VideoMM.