GleanVID:面向高效视频大语言模型的互补性令牌选择
GleanVID: Complementary Token Selection for Efficient Video Large Language Models
浏览论文内容
中文总结 AI 辅助
提出无需训练的GleanVID框架,通过跨帧互补性令牌选择,在仅用25%视觉令牌时保留Qwen3-VL 98.6%性能并降低44.7%延迟,实现高效视频理解。
中文摘要 AI 辅助
视频大语言模型(VideoLLMs)已实现强大的视频理解能力,但由于大量视觉令牌的存在,产生了巨大的推理开销。现有VideoLLM令牌压缩方法主要依赖与选择无关的评分,忽视了跨帧互补性,从而在帧间保留了冗余证据。相反,我们将视频令牌选择视为一个渐进式证据累积过程。其目标是在有限的令牌预算下,保留既在个体上具有信息量、又在整体上具有互补性的视觉证据。基于这一见解,我们提出了GleanVID,一个无需训练的视频大语言模型推理加速框架。具体而言,GleanVID首先根据时间新颖性在帧间分配全局令牌预算,然后通过联合考虑局部代表性和子空间互补性来选择令牌,从而保留更丰富且冗余更少的视觉证据。在多种VideoLLM和基准上的大量实验表明,GleanVID持续实现了最先进的性能。值得注意的是,仅使用25%的视觉令牌,GleanVID保留了Qwen3-VL原始性能的98.6%,同时将其预填充延迟降低了44.7%。在LLaVA-OV-7B上,GleanVID在25%保留率下的性能甚至略微超过了原始模型。
英文摘要
Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and consequently retaining redundant evidence across frames. Instead, we view video token selection as a progressive evidence accumulation process. It aims to retain visual evidence that is individually informative and collectively complementary under a limited token budget. Building on this insight, we introduce GleanVID, a training-free inference acceleration framework for VideoLLMs. Specifically, GleanVID first allocates the global token budget across frames according to temporal novelty and then selects tokens by jointly considering local representativeness and subspace complementarity, thereby preserving richer and less redundant visual evidence. Extensive experiments across diverse VideoLLMs and benchmarks demonstrate that GleanVID consistently achieves state-of-the-art performance. Notably, with only 25% of visual tokens, GleanVID preserves 98.6% of Qwen3-VL's original performance while reducing its prefill latency by 44.7%. On LLaVA-OV-7B, GleanVID at a 25% retention ratio even slightly surpasses the original model.
发表机构
- Beihang University(北京航空航天大学)
- Communication University of China(中国传媒大学)
机构由 AI 辅助整理,请以论文原文为准。