AI 中文总结
该研究针对VLMs的语义漂移问题,提出无训练的CaRe框架,通过两个互补模块校准视觉表示,在剪枝94.4%视觉令牌时保留96.4%原性能,使推理速度提升最高2.30倍。
AI 中文摘要
大型视觉语言模型(VLMs)因视觉令牌序列过长而面临难以承受的推理开销。现有视觉令牌缩减方法主要通过剪枝或压缩冗余令牌来提升效率,但未验证生成的表示是否与原始表示保持语义一致性。将原始N个令牌的视觉序列映射为K个令牌可能会丢弃、稀释或错误分配关键视觉线索,引发严重的语义漂移,导致VLM的理解出现偏差。本文首先提出视觉令牌缩减的“推理前校准”原则,并设计了CaRe这一无训练鲁棒框架,该框架在推理前校准紧凑的视觉表示以保留VLMs的语义保真度。CaRe由两个互补模块组成:1)抗扰动校准锚定,该模块在多方向扰动下识别具有稳定模型侧影响的校准锚;2)置信门控令牌校准,该模块从未被选中的令牌中提取可靠的校准信号并注入锚中。在不同VLM架构和基准上的大量评估验证了CaRe优于最先进的令牌缩减基线。在剪枝94.4%的视觉令牌的同时,我们的方法保留了原始全令牌性能的96.4%,与未剪枝的普通模型相比,端到端推理速度最高提升了2.30倍。
英文摘要
Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token reduction methods mainly improve efficiency by pruning or compressing redundant tokens without examining whether the resulting representation remains semantically consistent with the original representation. Mapping the original N-token visual sequence to K tokens may discard, dilute, or misassign critical visual cues, triggering severe semantic drift that deviates the VLM's understanding. In this paper, we first introduce the principle of 'Calibrate Before Reason' to visual token reduction and propose CaRe, a training-free robust framework that calibrates compact visual representations before reasoning to preserve semantic fidelity in VLMs. CaRe consists of two mutually complementary modules: 1) Perturbation-Robust Calibration Anchoring, which identifies calibration anchors with stable model-side influence under multi-directional perturbations; 2) Confidence-Gated Token Calibration, which extracts reliable calibration signals from unselected tokens and injects them into anchors. Extensive evaluations across diverse VLM architectures and benchmarks verify that CaRe outperforms state-of-the-art token reduction baselines. While pruning 94.4% of visual tokens, our method retains 96.4% of the original full-token performance, delivering up to 2.30 times faster end-to-end inference speed relative to unpruned vanilla models.