AI 中文总结
本研究针对视觉-语言模型中视觉token导致的推理瓶颈,提出无需训练的DIVE框架,通过动态迭代构建视觉证据,在减少88.9%视觉token的同时保留98.2%的平均性能。
AI 中文摘要
视觉-语言模型(VLMs)中的视觉输入通常被编码为比文本长得多的token序列,这使得视觉token成为高效推理的主要瓶颈。近期大量方法通过对token重要性打分并在单次前向传播中剪去低分token来解决该瓶颈。然而,单次打分是不够的,因为token与提示相关的有用性取决于已保留的证据。受此见解启发,我们提出DIVE(Dynamic Iterative Visual Evidence Construction,动态迭代视觉证据构建),这是一个无需训练的框架,它将视觉token剪重视为动态证据构建。DIVE反复选择具有最高残差条件得分的剩余token,更新视觉和提示残差以抵消已解释的证据,并重新评估剩余token。这种选择-更新-重新评估的过程构建了一组互补的、与提示相关的保留证据。在8个图像理解基准上的实验表明,DIVE在不同token预算下均能稳定保持性能。在视觉token减少88.9%的情况下,DIVE保留了未压缩模型平均性能的98.2%。代码可在该https URL获取。
英文摘要
Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottleneck for efficient inference. Abundant recent methods address this bottleneck by scoring token importance and pruning low-scoring tokens in a single pass. However, one-shot scoring is insufficient because a token's prompt-relevant usefulness depends on the evidence already retained. Motivated by this insight, we introduce DIVE (Dynamic Iterative Visual Evidence Construction), a training-free framework that recasts visual-token pruning as dynamic evidence construction. DIVE repeatedly selects the remaining token with the highest residual-conditioned score, updates the visual and prompt residuals to discount the evidence already explained, and re-evaluates the remaining tokens. This select-update-re-evaluate process builds a retained set of complementary, prompt-relevant evidence. Experiments across eight image-understanding benchmarks show that DIVE consistently preserves performance across token budgets. With an 88.9% reduction in visual tokens, DIVE retains 98.2% of the uncompressed model's average performance. Code is available at https://github.com/Zhong-Chenchen/DIVE.git.