arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01224cs.CV

S$^2$Prune:面向多模态大语言模型的空间结构化视觉令牌剪枝

S$^2$Prune: Spatially Structured Visual Token Pruning for Multimodal Large Language Models

  • College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息与电子工程学院)

机构由 AI 辅助整理,请以论文原文为准。

Yuanyuan Jia, Shunpu Tang, Qianqian Yang

AI总结:

S$^2$Prune是一种无需训练的多模态大语言模型视觉令牌剪枝方法,通过空间结构化策略保留图像覆盖、适配局部结构,在Qwen2.5-VL-7B-Instruct上仅用少量令牌即可保持较高性能。

AI中文摘要:

视觉令牌剪枝通过仅保留部分视觉令牌来降低多模态大语言模型(MLLM)的推理开销。现有方法通常基于重要性或冗余度选择令牌,但我们观察到这些标准会在不同输入间产生稳定的空间偏差,且并不总能优于简单的均匀网格采样,这凸显了广泛空间覆盖的价值。受此启发,我们提出S$^2$Prune,一种无需训练的剪枝方法,在保留空间覆盖的同时,使令牌密度适配局部图像结构。我们首先将图像划分为区域,为每个区域分配至少一个令牌以保留覆盖范围;剩余令牌预算则根据拉普拉斯变化进行分配,为结构更丰富的区域提供更多令牌。接着,我们利用首个解码器块计算的早期表征变化(ERC),在每个区域内选择代表性令牌。我们在多种设置和两种MLLM架构上对S$^2$Prune进行评估,在Qwen2.5-VL-7B-Instruct上,它在所有评估的无需训练的剪枝方法中取得了最高的平均准确率;仅使用原始576个视觉令牌中的32个时,仍保留了全模型79.3%的性能。代码可在该https URL获取。

英文摘要:

Visual token pruning reduces the inference overhead of multimodal large language models (MLLMs) by retaining only a subset of visual tokens. Existing methods usually select tokens based on importance or redundancy. However, we observe that these criteria produce stable spatial biases across inputs and do not always outperform simple Uniform Grid sampling, highlighting the value of broad spatial coverage. Motivated by this, we propose S$^2$Prune, a training-free pruning method that preserves spatial coverage while adapting token density to local image structure. We first divide the image into regions and assign at least one token to each region to preserve coverage. The remaining token budget is then distributed according to Laplacian variation, giving more tokens to regions with richer structure. We then use Early Representation Change (ERC), computed from the first decoder block, to select representative tokens within each region. We evaluate S$^2$Prune across diverse settings and two MLLM architectures. On Qwen2.5-VL-7B-Instruct, it achieves the highest average accuracy among the evaluated training-free pruning methods. With only 32 of the original 576 visual tokens, it still retains 79.3% of the full-model performance. Code is available at https://github.com/yuanyuanjia71-spec/S2Prune.

补充信息

↑