arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03158cs.CVcs.CLcs.LG

谁为剪枝后的内容发声?视觉令牌剪枝作为覆盖优化

Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization

Qingchan Zhu, Weihang You, Hanqi Jiang, Changdi Yang, Tianming Liu, Geng Yuan

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对视觉语言模型剪枝仅关注保留令牌的缺陷,提出无需训练的CoverPruner,将剪枝建模为表征覆盖最大化,在多架构及压缩率下实现最优平均准确率,激进压缩时增益最大。

中文摘要 AI 辅助

视觉令牌剪枝可降低视觉语言模型(VLMs)的推理成本,但多数方法仅关注保留哪些令牌。这种保留令牌的视角会保留冗余的高分令牌,同时被丢弃的证据缺乏相近的代表。本文提出CoverPruner,一种无需训练的剪枝器,它关注互补的需求侧问题:移除某令牌后,哪个留存的原始令牌会为目标VLM代表它。CoverPruner将剪枝建模为表征覆盖最大化(RCM),以查询加权需求覆盖完整的投影视觉令牌集。它通过投影空间覆盖和轻量的第一层注意力探针实例化RCM。在多个VLM架构和压缩率下,CoverPruner在所有对比方法中实现最优平均准确率,且在激进压缩下通常获得最大增益。

英文摘要

Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.

发表机构

  • School of Computing, University of Georgia(佐治亚大学计算学院)
  • College of Engineering, Northeastern University(东北大学工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑