arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03389cs.CVcs.LG

从修补到剪枝:视觉语言模型中的视觉计算

From Patching to Pruning Visual Computation in Vision Language Models

Rahul Chowdhury, Timothy A Rupprecht, Xuan Shen, Shaoyi Huang, Pu Zhao, Yanzhi Wang

首次发表
浏览论文内容

中文总结 AI 辅助

P2P框架通过将激活修补转化为推理时计算旁路,在保留序列结构的同时剪枝视觉计算,实现3%容差下94%精度和55%FLOPs减少,并揭示VLM视觉处理在深度上的非均匀分布。

中文摘要 AI 辅助

视觉语言模型(VLMs)因每个视觉标记都要经过每个解码器层的注意力机制和MLP投影处理而产生大量推理成本,即使在许多深度层上,针对特定标记的视觉计算并非必要。我们引入了受机制可解释性启发的Patch-to-Prune(P2P)框架,这是一个无需训练的方法,将激活修补从诊断工具转变为推理时的计算旁路。P2P执行验证引导的前向和后向层扫描,以识别那些在用户指定的精度容差内,其视觉标记投影输出可以被固定的中性代理激活向量替换的解码器区域。与传统的标记剪枝方法不同,P2P保留了序列长度、标记顺序、位置信息、注意力掩码和残差路径,从而在不移除标记或修改预训练模型权重的情况下剪枝计算。我们在来自Qwen2.5-VL和LLaVA家族的四个VLM上,使用互不相交的校准、验证和测试分区,在七个多模态基准上评估了P2P。在3%容差下,P2P保留了约94%的密集精度,同时将FLOPs减少了55%。除了这些效率提升,我们的逐层分析表明,VLM中的视觉处理在解码器深度上非均匀分布:早期和晚期层通常需要很少的特定标记视觉计算,而中间层似乎执行了大部分与任务相关的视觉整合,使得后续推理在很大程度上依赖于已经嵌入在共享残差和文本表示中的视觉信息。这使得P2P既是一个高效的推理框架,也是理解VLM中视觉信息处理的因果透镜。

英文摘要

Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass. P2P performs validation-guided forward and backward layer sweeps to identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance. Unlike conventional token-pruning methods, P2P preserves the sequence length, token order, positional information, attention mask, and residual pathways, thereby pruning computation without removing tokens or modifying the pretrained model weights. We evaluate P2P on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks using mutually disjoint calibration, validation, and test partitions. P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%. Beyond these efficiency gains, our layer-wise analysis suggests that visual processing in VLMs is non-uniformly distributed across decoder depth: early and late layers often require little token-specific visual computation, whereas intermediate layers appear to perform most task-relevant visual integration, enabling later reasoning to rely largely on visual information already embedded in shared residual and textual representations. This makes P2P both an efficient inference framework and a causal lens into visual information processing in VLMs.

发表机构

  • Northeastern University(东北大学)
  • EmbodyX Inc.(EmbodyX公司)
  • Zhejiang University(浙江大学)
  • Stevens Institute of Technology(史蒂文斯理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑