发表机构
Nanjing University; Geely Automobile Research Institute (Ningbo) Co., Ltd.; Goertek(南京大学; 吉利汽车研究院(宁波)有限公司; 歌尔股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TReVS提出无需训练的视觉令牌剪枝框架,融合文本相关性与视觉显著性,利用高方差注意力头,在LLaVA-1.5-7B上剪枝94.4%视觉令牌并保留92.8%性能。
AI 中文摘要
视觉语言模型(VLMs)在视觉理解和推理方面表现出色,但由于视觉令牌数量庞大,常常产生高昂的推理成本。最近的视觉令牌剪枝方法越来越多地采用两阶段范式:首先在视觉编码器之后移除视觉冗余令牌,然后在大型语言模型(LLM)内丢弃与文本查询无关的令牌。然而,由于第一阶段通常仅依赖视觉编码器的显著性,它可能过早地消除与查询相关的令牌,从而剥夺后续文本引导阶段的关键视觉证据。我们的实证分析表明,在第一阶段剪枝中纳入查询引导能更好地保留任务相关证据,并且始终优于仅基于视觉显著性的剪枝。我们进一步发现,高方差注意力头对文本查询更为敏感,并为第二阶段剪枝产生更具区分性的文本到视觉注意力信号。受这些发现的启发,我们提出了TReVS,一个无需训练框架,将文本相关性与视觉编码器显著性相结合用于LLM前剪枝,并利用高方差注意力头在LLM的浅层到中间层移除任务无关令牌。在LLaVA-1.5-7B上,TReVS在剪枝94.4%的视觉令牌的同时,保留了未剪枝基线性能的92.8%,优于先前的最先进方法。
英文摘要
Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically relies solely on vision-encoder saliency, it may prematurely eliminate query-relevant tokens, depriving the subsequent text-guided stage of critical visual evidence. Our empirical analysis shows that incorporating query guidance into first-stage pruning better preserves task-relevant evidence and consistently improves performance over vision-only saliency-based pruning. We further find that high-variance attention heads are more sensitive to the textual query and yield more discriminative text-to-vision attention signals for second-stage pruning. Motivated by these findings, we propose TReVS, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM. On LLaVA-1.5-7B, TReVS retains 92.8% of the unpruned baseline performance while pruning 94.4% of visual tokens, outperforming prior state-of-the-art methods.