SPIDER:多模态大语言模型中的多层语义令牌剪枝与自适应子层跳过
SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models
浏览论文内容
中文总结 AI 辅助
SPIDER提出无需训练的多层语义令牌剪枝与自适应子层跳过机制,解决多模态大模型数据与计算双重冗余,在LLaVA-NeXT-7B上减79% FLOPs且保持96%性能。
中文摘要 AI 辅助
多模态大语言模型面临显著的计算效率挑战,这些挑战源于两个不同但又相互关联的方面:数据冗余和计算冗余。尽管大多数方法通过从视觉编码器的输出中剪枝视觉令牌来处理数据冗余,或利用分块重要性在LLM解码器中计算冗余,但层间更细粒度的表示偏移以及层内部的分布差异尚未被充分探索。在本工作中,我们全面研究了这种双重层面的低效问题。我们假设,视觉编码器的中间层令牌应被考虑用于有效的视觉令牌剪枝,因为语义焦点在层间发生转移,中间层令牌捕获了更多详细的、以对象为中心的信息,而这些信息可能被更深层所抽象化。此外,我们揭示了注意力机制和前馈网络在不同LLM解码器层中的差异性贡献。基于这些发现,我们提出了SPIDER,一个无需训练框架,它集成了多层语义视觉令牌剪枝与自适应子层跳过机制。实验评估表明,SPIDER在各种多模态大语言模型架构和缩减比率下均能保持强劲性能。例如,在LLaVA-NeXT-7B上,SPIDER将FLOPs减少了79%,同时保持了基线性能的96%。
英文摘要
Multimodal Large Language Models face significant efficiency challenges that stem from two distinct yet coupled sources: data redundancy and computational redundancy. While most methods focus on data redundancy by pruning visual tokens from the output of the visual encoder or computing redundancy in LLM decoders using blockwise importance, the finer-grained inter-layer representation shifts and the distribution differences within the layers themselves have not been fully explored. In this work, we comprehensively investigate this dual-level inefficiency. We posit that intermediate layer tokens from vision encoders should be considered for effective visual token pruning, as semantic focus shifts across layers, with middle-layer tokens capturing more detailed object-centric information that deeper layers may abstract away. Furthermore, we reveal the differential contributions of Attention and FFNs across distinct LLM decoder layers. Building upon these discoveries, we propose \textbf{SPIDER}, a training-free framework that integrates multi-layer \underline{\textbf{S}}emantic visual token \underline{\textbf{P}}run\underline{\textbf{I}}ng with an a\underline{\textbf{D}}aptive sub-lay\underline{\textbf{ER}} skipping mechanism. Experimental evaluations demonstrate that SPIDER consistently maintains strong performance across various MLLM architectures and reduction ratios. For instance, on LLaVA-NeXT-7B, SPIDER reduces FLOPs by $79\%$ while maintaining 96$\%$ of the baseline performance.
发表机构
- Alibaba Cloud(阿里云)
- University of Science and Technology of China(中国科学技术大学)
- Maynooth University(梅努斯大学)
- Fudan University(复旦大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。