arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SepPrune:一种基于分隔符的高效多模态大语言模型剪枝框架

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models

Yuchen Wang, Qihui Zhu, Yang Liu, Xiaoyan Sun, Siying Wu

arXiv 2607.25818首次发表:更新:

发表机构

University of Science and Technology of China; ChangXin Memory Technologies(中国科学技术大学; 长鑫存储技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型视觉令牌计算成本高的问题,提出SepPrune剪枝方法,以模态分隔符令牌为统一查询来排序和选择视觉令牌,复用内置投影参数,无需架构更改,在Qwen2.5-VL-7B上实验取得最优性能。

AI 中文摘要

近期的多模态大语言模型,如Qwen2.5-VL和InternVL3,对高分辨率输入会生成大量视觉令牌,导致计算成本高昂。现有视觉令牌剪枝方法要么依赖跨模态注意力且无法在预填充阶段前剪枝,要么依赖高计算开销的多样性估计。我们发现视觉和文本令牌的注意力分数在模态分隔符令牌处达到峰值,基于此提出SepPrune。它是一种高效、无需训练、即插即用的剪枝方法,用分隔符令牌作为统一查询来排序和选择信息丰富的视觉令牌。SepPrune复用大语言模型的内置投影参数,无需架构更改。在Qwen2.5-VL-7B上的实验表明,SepPrune实现了最优性能,去除80.2%的视觉令牌时保留了96.3%的原始准确率。

英文摘要

Recent multimodal large language models (MLLMs), such as Qwen2.5-VL and InternVL3, generate large numbers of vision tokens for high-resolution inputs, leading to substantial computational cost. Existing vision token pruning methods either depend on cross-modal attention and cannot prune before the prefill stage, or rely on diversity estimation with high computational overhead. We observe that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities. Based on this observation, we propose SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informative vision tokens. SepPrune reuses the LLM's built-in projection parameters and requires no architectural changes. Experiments on Qwen2.5-VL-7B show that SepPrune achieves state-of-the-art performance, retaining 96.3% of the original accuracy while removing 80.2% of vision tokens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑