SinkPruner:面向多模态大语言模型的无汇视觉令牌剪枝
SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究针对多模态大语言模型处理长视觉令牌序列时计算开销大的问题,提出无训练的SinkPruner剪枝框架,通过视觉净化器和文本引导剪枝器,在12个图像-语言和4个视频-语言基准上实现高效推理,在令牌减少89%时保留了LLaVA-1.5和Qwen2.5-VL的大部分性能。
中文摘要 AI 辅助
尽管多模态大语言模型(MLLM)具备强大的多模态理解能力,但在处理长视觉令牌序列时会产生大量计算开销。为降低推理成本,近期研究探索了基于视觉中心或文本引导策略的视觉令牌剪枝方法。然而,这些方法常忽略高范数离群令牌,即特征范数异常大的令牌,导致剪枝决策欠佳。本研究表明,此类高范数离群令牌在特征和空间维度均高度冗余,却常被现有方法误判为有效线索。基于此观察,我们提出SinkPruner,一种用于高效MLLM推理的无训练视觉令牌剪枝框架。SinkPruner采用由粗到精的设计,包含两个关键模块:视觉净化器,用于过滤高范数冗余并缓解注意力汇和注意力分散;文本引导剪枝器,用于进一步保留与文本查询语义对齐的令牌。在12个图像-语言和4个视频-语言基准上的大量实验验证了该框架的有效性、效率和泛化性。值得注意的是,在令牌减少89%的情况下,SinkPruner保留了LLaVA-1.5(Qwen2.5-VL)96.5%(91.8%)的原始性能。实验进一步表明,我们的视觉净化器在提升现有剪枝方法性能方面展现出良好的可迁移性。我们的代码可在this https URL获取。
英文摘要
Despite their strong multimodal understanding ability, multimodal large language models (MLLMs) incur substantial computational overhead when processing long visual token sequences. To reduce inference costs, recent studies have explored visual token pruning through vision-centric or text-guided strategies. However, these methods often overlook high-norm outlier tokens, i.e., tokens with abnormally large feature norms, leading to suboptimal pruning decisions. In this work, we show that such high-norm outlier tokens are highly redundant in both feature and spatial dimensions, yet are often mistakenly preserved as informative cues by existing methods. Motivated by this observation, we propose SinkPruner, a training-free visual token pruning framework for efficient MLLM inference. SinkPruner follows a coarse-to-fine design with two key modules: a visual sanitizer that filters high-norm redundancies and alleviates attention sink and attention dispersion, and a text-guided pruner that further retains tokens semantically aligned with the text query. Extensive experiments on twelve image-language and four video-language benchmarks demonstrate the effectiveness, efficiency, and generalizability of our framework. Notably, SinkPruner preserves 96.5% (91.8%) of the original performance of LLaVA-1.5 (Qwen2.5-VL) under an 89% token reduction. Experiments further indicate that our visual sanitizer exhibits promising transferability in enhancing the performance of existing pruning methods. Our code is available at https://github.com/LaVi-Lab/SinkPruner.
发表机构
- Weitu AI(微图AI)
- Peking University(北京大学)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。