发表机构
Harbin Institute of Technology; Shenzhen Loop Area Institute(哈尔滨工业大学; 深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大语言模型视觉标记冗余导致的推理延迟问题,本文分析标记表示动力学,提出无需训练的MSDG-Prune剪枝方法,利用更新幅度与方向保留显著多样信息,在LLaVA-NeXT上以5.6%标记保留91.9%性能并实现7.8倍加速。
AI 中文摘要
多模态大语言模型(MLLMs)因长视觉标记序列而产生高推理延迟。现有的剪枝方法通常使用注意力图或输出特征来估计标记的重要性或冗余性。最近的几种方法也利用表示变化,但这些变化何时以及如何反映前景显著性和语义一致性仍未得到充分理解。我们分析了视觉标记表示在编码器深度上的动力学,并发现了两个结论。首先,标记更新幅度与前景显著性之间的关系是层相关的:在两个深度区间内,大的标记更新集中在前景区域,这些区间被中间深度处的若干汇(sink)主导层分隔。其次,标记更新方向之间的相似性比编码器输出特征之间的相似性更能区分同类与不同类标记。基于这些发现,我们提出了MSDG-Prune,一种无需训练的方法,利用更新幅度和方向来保留显著且多样的视觉信息。具体来说,我们根据更新方向相似性对标记进行分组,并使用在选定深度窗口上从更新幅度导出的查询加权显著性进行分组级标记剪枝。在四个MLLM上的大量实验证明了MSDG-Prune的有效性和泛化性。在LLaVA-NeXT上,它仅使用5.6%的视觉标记就保留了平均91.9%的未压缩性能,同时实现了7.8倍的预填充加速。代码可在该https URL获取。
英文摘要
Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several recent approaches also exploit representation changes, but when and how these changes reflect foreground saliency and semantic consistency remain insufficiently understood. We analyze visual token representation dynamics across encoder depth and uncover two findings. First, the relationship between token update magnitudes and foreground saliency is layer-dependent: large token updates concentrate on foreground regions in two depth intervals, separated by several sink-dominated layers at intermediate depths. Second, similarities between token update directions better distinguish same-class from different-class tokens than those between encoder output features. Building on these findings, we propose MSDG-Prune, a training-free method that uses update magnitudes and directions to preserve salient and diverse visual information. Specifically, we group tokens by update-direction similarity and use query-weighted saliency derived from update magnitudes across a chosen depth window for group-wise token pruning. Extensive experiments across four MLLMs demonstrate the effectiveness and generalizability of MSDG-Prune. On LLaVA-NeXT, it retains 91.9% of uncompressed performance on average with only 5.6% of visual tokens, while achieving a 7.8x prefilling speedup. Code is available at https://github.com/liweixuan-hitsz/MSDG-Prune.
CommentsPreprint. 33 pages, 17 figures, 19 tables