arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34861cs.CVcs.LG

当文本重要时:视觉语言模型中视觉令牌剪枝的设计原则

When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model

Minchan Kang, Kyeonghye Park, Seoyoung Cho, Daeshik Kim, Yucheol Cho

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种无需训练的视觉令牌剪枝方法,通过分离早期视觉引导剪枝与延迟文本引导重新选择,在八个基准和三个模型上显著提升性能恢复,同时保持较低延迟。

中文摘要 AI 辅助

视觉令牌剪枝作为一种降低大型视觉语言模型计算成本的实用方法已被广泛研究。然而,该方法难以保留必要的视觉信息,这可能导致显著的性能下降。具体而言,基于图像的令牌选择可能忽略与任务相关的细节,而基于文本引导的令牌选择可能无法捕获复杂推理所需的文本-视觉关系。我们发现,过早应用文本引导会限制其识别与答案相关的视觉区域的能力,而文本到视觉的注意力在解码器的中间层变得更具信息性。这一发现促使我们提出一种无需训练的方法,该方法将早期的视觉引导剪枝与延迟的文本引导重新选择分离。我们首先使用视觉编码器注意力剪枝视觉令牌,保留额外候选令牌直到解码器中间点,然后使用文本到视觉注意力确定最终的视觉令牌集。在八个基准测试和三个模型上,我们的方法在80%和90%剪枝率下的性能恢复平均分别比最佳基线高出11.10和16.84个百分点,且与大多数基线相比,LLM预填充延迟相当或更低。源代码在此https URL公开提供。

英文摘要

Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection can overlook task-relevant details, while text-guided token selection may fail to capture the text--visual relationships needed for complex reasoning. We find that applying textual guidance too early can limit its ability to identify answer-relevant visual regions, whereas text-to-visual attention becomes more informative at intermediate decoder depths. This finding motivates our training-free method, which separates early vision-guided pruning from deferred text-guided reselection. We first prune visual tokens using vision-encoder attention, retain additional candidates until the decoder midpoint, and then use text-to-visual attention to determine the final visual-token set. Across eight benchmarks and three models, our method outperforms the best-performing baselines by an average of 11.10 and 16.84 percentage points in performance recovery at 80% and 90% pruning, respectively, with comparable or lower LLM-prefill latency than most baselines. The source code is publicly available at https://github.com/kmc3661/DeFT

发表机构

  • Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
  • Hanbat National University(韩巴国立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑