谁为剪枝后的内容发声?视觉令牌剪枝作为覆盖优化
Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization
浏览论文内容
中文总结 AI 辅助
本文针对视觉语言模型剪枝仅关注保留令牌的缺陷,提出无需训练的CoverPruner,将剪枝建模为表征覆盖最大化,在多架构及压缩率下实现最优平均准确率,激进压缩时增益最大。
中文摘要 AI 辅助
视觉令牌剪枝可降低视觉语言模型(VLMs)的推理成本,但多数方法仅关注保留哪些令牌。这种保留令牌的视角会保留冗余的高分令牌,同时被丢弃的证据缺乏相近的代表。本文提出CoverPruner,一种无需训练的剪枝器,它关注互补的需求侧问题:移除某令牌后,哪个留存的原始令牌会为目标VLM代表它。CoverPruner将剪枝建模为表征覆盖最大化(RCM),以查询加权需求覆盖完整的投影视觉令牌集。它通过投影空间覆盖和轻量的第一层注意力探针实例化RCM。在多个VLM架构和压缩率下,CoverPruner在所有对比方法中实现最优平均准确率,且在激进压缩下通常获得最大增益。
英文摘要
Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.
发表机构
- School of Computing, University of Georgia(佐治亚大学计算学院)
- College of Engineering, Northeastern University(东北大学工程学院)
机构由 AI 辅助整理,请以论文原文为准。