AI 中文总结
该研究针对视觉语言模型(VLM)海量视觉令牌导致推理成本高的问题,提出基于跨模态残差(CMR)的无训练压缩方法SIEVE,在LLaVA-NeXT-7B上仅保留11.1%视觉令牌,实现显著加速与缓存缩减且性能损失极小。
AI 中文摘要
丰富的视觉信息增强了视觉语言模型(VLM)的感知能力,但海量视觉令牌会提升推理成本。现有的视觉令牌剪枝方法依赖基于相似度的引导,该方法利用文本-视觉、视觉-视觉令牌的成对相关性进行压缩,然而这类方法仅能捕获局部层级信号,忽略了VLM的整体推理过程。在本文中,我们重新审视VLM的推理过程,提出一种可补充基于相似度引导的新型高效引导方案。具体而言,我们发现一个关键现象:随着大语言模型(LLM)层数加深,文本令牌会通过自注意力持续聚合视觉信息,并逐步将部分视觉内容吸收到文本表示中。为量化该现象,我们从几何表示视角提出跨模态吸收(Cross Modal Absorption, CMA),用于衡量文本吸收了多少视觉信息,结果表明更深层中更多视觉令牌可由文本子空间近似解释。据此,我们提出跨模态残差(Cross Modal Residual, CMR),通过Tikhonov正则化最小二乘法将视觉令牌投影到文本子空间,并利用重建残差量化无法由文本解释的视觉信息。最后,基于CMR,我们提出SIEVE,一种无需训练的视觉令牌压缩方法,结合CMR、文本注意力相关性与残差空间多样性,以保留任务相关且具有互补性的令牌。在多种VLM架构上的实验验证了SIEVE的有效性:例如在LLaVA-NeXT-7B上,SIEVE仅保留11.1%的视觉令牌,同时维持97.5%的原始平均性能,实现3.62倍的预填充加速、2.49倍的端到端加速以及6.02倍的KV缓存缩减。
英文摘要
Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on similarity-based guidance, which exploits pairwise text-vision and vision-vision token correlations for compression. However, such methods only capture local layer-level signals and overlook the whole inference process in VLM. In this paper, we revisit VLM inference and present a new efficient guidance scheme that complements similarity-based guidance. In particular, we identify a key observation: as LLM layers deepen, text tokens continuously aggregate visual information via self-attention and progressively absorb partial visual content into textual representations. To quantify this phenomenon, we propose Cross Modal Absorption (CMA) from a geometric representation perspective to measure how much visual information is absorbed by text, revealing that more visual tokens in deeper layers can be approximately explained by the text subspace. We accordingly propose Cross Modal Residual (CMR). It projects visual tokens onto the text subspace via Tikhonov regularized least squares and exploits reconstruction residuals to quantify visual information that cannot be explained by text. Finally, based on CMR, we present SIEVE, a training-free visual token compression method that combines CMR, text-attention relevance, and residual-space diversity to retain task-relevant and complementary tokens. Experiments on diverse VLM architectures verify the effectiveness of SIEVE. For instance, on LLaVA-NeXT-7B, SIEVE keeps only $11.1\%$ of visual tokens while preserving $97.5\%$ of the original average performance, achieving $3.62\times$ prefill speedup, $2.49\times$ end-to-end speedup, and a $6.02\times$ KV-cache reduction.