修剪前的中心化:用于多模态大语言模型(LVLMs)中基于多样性的视觉标记修剪的轻量级几何校正
Centering before Pruning: Lightweight Geometry Correction for Diversity-Based Visual Token Pruning in LVLMs
- Inha University(仁荷大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LVLMs视觉标记序列冗余导致推理成本高的问题,提出Cen-Prune方法,通过保留原始空间独特性并结合中心化余弦相似度,提升了现有多样性视觉标记修剪器的性能。
AI中文摘要:
多模态大语言模型(LVLMs)因视觉标记序列冗长且冗余度高,会产生高昂的推理成本。基于多样性的修剪通过成对余弦相似度选择标记子集来降低该成本,但我们发现原始视觉标记间的相似度强烈集中在正区间,限制了其区分非冗余标记的能力。提升这种区分度的自然方式是在计算余弦相似度前对标记特征进行中心化,中心化确实能揭示更丰富的成对结构,但单独使用时却意外降低了修剪性能。我们表明,这种明显的矛盾源于原始几何不仅代表成对多样性,还隐含偏向全局独特的标记,这些标记往往包含语义信息内容。中心化能更好地解决子集多样性,但会失去这种有用的标记级偏好,说明原始几何中多样性与独特性相互纠缠。基于此分析,我们提出中心化几何修剪器(Cen-Prune),该方法使用中心化余弦相似度测量子集多样性,同时保留原始空间的独特性作为互补的标记级偏好。这种轻量级即插即用校正未改变底层选择机制,计算开销可忽略不计。在多个图像和视频理解基准及LVLM架构上的大量实验表明,Cen-Prune在现有基于多样性的修剪器中,能为整体性能提供稳健提升。
英文摘要:
Large vision-language models (LVLMs) incur substantial inference costs due to their long and highly redundant visual-token sequences. Diversity-based pruning mitigates this cost by selecting token subsets based on pairwise cosine similarity. We find, however, that similarities between raw visual tokens are strongly concentrated in the positive range, limiting their ability to distinguish non-redundant tokens. A natural way to improve this resolution is to center token features before computing cosine similarity. Centering indeed reveals a substantially richer pairwise structure, yet unexpectedly degrades pruning performance when used alone. We show that this apparent contradiction arises because the raw geometry does more than represent pairwise diversity: it also implicitly favors globally distinctive tokens, which tend to contain semantically informative content. Centering better resolves subset diversity but loses this useful token-wise preference, revealing that diversity and distinctiveness are entangled in the raw geometry. Based on this analysis, we propose the \textbf{Cen}tered Geometry \textbf{Prune}r (Cen-Prune), which measures subset diversity using centered cosine similarity while retaining raw-space distinctiveness as a complementary token-wise preference. This lightweight, plug-and-play correction leaves the underlying selection mechanism unchanged and incurs negligible computational overhead. Extensive experiments across multiple image- and video-understanding benchmarks and LVLM architectures demonstrate that Cen-Prune provides robust improvements in overall performance across existing diversity-based pruners.