arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34558cs.CV

ACPruner:在LVLMs中将视觉令牌剪枝视为有偏注意力覆盖最大化

ACPruner: Visual Token Pruning as Biased Attention Coverage Maximization in LVLMs

Xu Li, Yuxuan Liang, Yi Zheng, Zhe Liu, Xiaolei Chen, Haotian Chen, Rui Zhu, Fan Shi, Xiangyang Xue

首次发表
浏览论文内容

中文总结 AI 辅助

ACPruner提出有偏注意力覆盖最大化方法,通过贪心选择紧凑视觉令牌子集,在多个LVLM骨干上实现高效推理与性能保持。

中文摘要 AI 辅助

大型视觉语言模型(LVLMs)因大量视觉令牌而面临显著的计算效率低下问题。现有的视觉令牌剪枝方法主要侧重于保留单独重要的令牌或选择相互多样的令牌。在这项工作中,我们从覆盖的角度重新审视视觉令牌剪枝,并将其表述为一个有偏的注意力覆盖最大化问题。关键思想是选择一个紧凑的令牌子集,其编码器侧的外向注意力能够共同覆盖图像,同时为信息更丰富的区域分配更高的覆盖优先级。基于这一视角,我们提出了ACPruner,一个无需训练的视觉令牌剪枝框架,用于高效的LVLM推理。ACPruner首先通过结合模态内显著性和模态间相关性来估计令牌重要性,然后从视觉编码器内的注意力模式中推导出令牌级覆盖,最后执行贪心选择以最大化所提出的覆盖目标。在多个LVLM骨干网络上的大量实验,包括LLaVA-1.5-7B/13B、LLaVA-NeXT-7B/13B、Qwen2.5-VL-7B和LLaVA-OneVision-7B,表明ACPruner在实现显著的端到端推理加速的同时,始终保持着强大的性能保持能力。

英文摘要

Large Vision-Language Models (LVLMs) face significant computational inefficiencies caused by the large number of visual tokens. Existing visual token pruning methods mainly focus on either retaining individually important tokens or selecting mutually diverse ones. In this work, we revisit visual token pruning from a coverage perspective and formulate it as a biased attention coverage maximization problem. The key idea is to select a compact token subset whose encoder-side outgoing attention can jointly cover the image while assigning higher coverage priority to more informative regions. From this perspective, we propose ACPruner, a training-free visual token pruning framework for efficient LVLM inference. ACPruner first estimates token importance by combining intra-modal saliency and inter-modal relevance, then derives token-wise coverage from attention patterns within the vision encoder, and finally performs greedy selection to maximize the proposed coverage objective. Extensive experiments across multiple LVLM backbones, including LLaVA-1.5-7B/13B, LLaVA-NeXT-7B/13B, Qwen2.5-VL-7B, and LLaVA-OneVision-7B, show that ACPruner consistently achieves strong performance retention while delivering substantial end-to-end inference speedups.

发表机构

  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑