发表机构
Maum AI Inc.(Maum AI公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对视觉-语言模型(VLMs)的视觉标记剪枝问题,通过EmbedLens分析标记角色,发现代表性剪枝方法存在标记角色偏差但与下游性能无直接关联,优化角色分配后表明保留非活跃标记可维持或提升性能,为剪枝方法优化提供了新视角。
AI 中文摘要
视觉-语言模型(VLMs)将图像处理为一系列视觉标记,这在推理过程中造成了巨大的计算瓶颈。近期的视觉标记剪枝方法通过移除看似冗余的标记来解决该问题,但目前尚不清楚这些剪枝决策与视觉标记的功能角色存在何种关联。本研究通过EmbedLens识别的标记角色视角分析视觉标记剪枝。我们首先表明,代表性剪枝方法表现出明显的标记角色偏差,但这些偏差与下游性能无直接关联。为更好理解该行为,我们优化了标记角色分配流程,并评估了角色保护型剪枝变体。结果显示,保留非活跃(non-alive)标记有时可维持或提升性能,这表明与直接语义对齐较弱的标记在剪枝场景下仍会影响模型行为。我们的代码公开于此https URL。
英文摘要
Vision-language models (VLMs) process an image as a sequence of visual tokens, which creates a substantial computational bottleneck during inference. Recent visual token pruning methods address this issue by removing seemingly redundant tokens, yet it remains unclear how these pruning decisions relate to the functional roles of visual tokens. In this work, we analyze visual token pruning through the lens of token roles identified by EmbedLens. We first show that representative pruning methods exhibit distinct token-role biases, but these biases do not directly correlate with downstream performance. To better understand this behavior, we refine the token-role assignment procedure and evaluate role-protected pruning variants. Our results show that preserving non-alive tokens can sometimes maintain or improve performance, suggesting that tokens with weak direct semantic alignment may still affect model behavior under pruning. Our code is publicly available at https://github.com/jaykim9870/Not_All_Redundant_Tokens_Are_Alike.
CommentsAccepted to ECCV 2026 workshop, UniWorld