DIPrune:基于双重重要性的任务感知令牌剪枝,用于高效多模态语言模型
DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models
浏览论文内容
中文总结 AI 辅助
针对多模态大语言模型无训练剪枝中的语义退化问题,提出DIPrune,通过双重重要性评分机制优化层内与层间信号,在LLaVA和Qwen-VL上取得最先进结果。
中文摘要 AI 辅助
最近针对多模态大语言模型(MLLMs)的无训练剪枝方法通过利用视觉冗余或文本-视觉注意力有效降低了计算开销。然而,由于其任务无关的设计或不可靠的注意力估计,这些方法经常遭受语义退化。基于我们的实证分析,我们发现这个问题源于浅层中的显著令牌通过数值惯性持续抑制新兴语义令牌,导致对深层推理至关重要的信号被过早丢弃。为了解决上述问题,从任务导向的角度出发,我们首先将无训练剪枝重新表述为对最终任务损失失真的最小化,并推导出一个可处理的、逐令牌的上界作为替代目标。具体而言,这一表述内在揭示了一个先前被忽视的层间项,该项考虑了跨层的梯度。相应地,在实现上,我们提出了DIPrune,一个基于排名的框架,采用双重重要性评分机制,联合优化层内静态特征显著性和层间动态语义演化。在LLaVA和Qwen-VL上的大量实验表明,DIPrune始终达到最先进的结果。
英文摘要
Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tractable, token-wise upper bound to serve as a surrogate objective. Specifically, this formulation inherently reveals a previously neglected inter-layer term that accounts for gradients across layers. Accordingly, for the implementation, we propose DIPrune, a rank-based framework that employs a dual importance scoring mechanism to jointly optimize intra-layer static feature saliency and inter-layer dynamic semantic evolution. Extensive experiments on LLaVA and Qwen-VL demonstrate that DIPrune consistently achieves state-of-the-art results.
发表机构
- Communication University of China(中国传媒大学)
- Beihang University(北京航空航天大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。