arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SlimVLM:面向高效视觉-语言模型的感知敏感性动态结构化剪枝与自适应视觉token选择

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models

Yaozhi Wen, Jialong Guo, Zhenliang Ni, Han Shu, Xinghao Chen

arXiv 2608.03580首次发表:更新:

AI 中文总结

SlimVLM是一种结构化剪枝框架,通过自适应视觉token选择和感知敏感性动态剪枝机制,在压缩视觉-语言模型的同时保持性能,在多模态基准测试中达到最优。

AI 中文摘要

视觉-语言模型(VLMs)在处理和理解文本与图像方面展现出卓越性能,但其庞大的参数规模带来了显著的计算开销,限制了在资源受限设备上的部署。剪枝技术在压缩大语言模型(LLMs)方面效果显著,但直接应用于VLMs会导致性能大幅下降,主要原因是冗余视觉token会干扰重要性估计。为此,我们提出SlimVLM,这是一种旨在压缩VLMs同时保留任务性能的结构化剪枝框架。我们为VLMs引入了自适应视觉token选择策略,该策略利用文本到视觉的平均注意力分数评估视觉token的重要性,基于设定阈值在剪枝过程中移除冗余token,从而优化重要性计算。考虑到不同模块对稀疏性的容忍度存在差异,我们还提出了感知敏感性动态剪枝机制,通过计算剪枝后与未剪枝模块输出间的线性重构误差,为每个模块确定合适的剪枝比例,确保整体性能稳定。实验结果表明,SlimVLM在多个多模态基准测试中均优于现有方法,取得了当前最优性能。

英文摘要

While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large parameter sizes lead to significant computational overhead, limiting their deployment on resource-constrained devices. While pruning has been effective for compressing Large Language Models (LLMs), directly applying it to VLMs leads to significant performance drops, largely due to redundant visual tokens interfering with importance estimation. To this end, we propose SlimVLM, a structured pruning framework designed to compress VLMs while preserving their task performance. We introduce an adaptive visual token selection strategy for VLMs that leverages average text-to-visual attention scores to assess the importance of visual tokens, removing redundant ones during pruning based on a set threshold, thereby optimizing the importance calculation. Recognizing the varying tolerance to sparsity across different modules, we also propose a Sensitivity-aware dynamic pruning mechanism that determines the appropriate pruning ratio for each module by calculating the linear reconstruction error between the outputs of the pruned and unpruned modules, ensuring overall performance stability. Experimental results show that SlimVLM outperforms existing methods across multiple multimodal benchmarks, achieving state-of-the-art performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑