从浅到深:在大型视觉-语言模型中使令牌剪枝与阶段角色对齐
Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs
浏览论文内容
中文总结 AI 辅助
针对LVLMs中视觉令牌冗余导致的计算开销,提出STD分层剪枝框架,依据网络各阶段角色自适应选择剪枝策略,在保持性能的同时显著减少令牌并加速推理。
中文摘要 AI 辅助
大型视觉-语言模型(LVLMs)因冗余的视觉令牌而产生高昂的计算成本。尽管在视觉编码器阶段,基于注意力的免训练多层剪枝已被探索为一种有效策略,但我们发现,在浅层进行剪枝会持续降低性能。本文旨在理解这一问题并寻求解决方案。通过分析跨网络深度的注意力模式,我们发现浅层主要充当边缘检测器,其注意力图混乱,而深层则经历局部主体识别和不稳定的语义聚合。为了解决剪枝策略与网络阶段之间的错位问题,我们提出了STD,一种分层令牌剪枝框架,它根据每个网络阶段的功能角色调整令牌选择机制。STD在浅层采用高频频谱分析,以确定性地保留结构边缘;在中间层使用高斯平滑注意力以维持空间连贯性;并在深层引入稳定性自适应触发器,仅在语义稳定阶段执行剪枝。大量实验表明,STD在LLaVA-1.5-7B上以88.9%的令牌缩减率超越最先进的剪枝方法1.1%,同时即插即用,与其他方法结合时也非常有效;在LLaVA-NeXT-7B上以94.4%的缩减率提升2.1%,在预填充阶段实现3.9倍加速。我们的代码将在该https URL发布。
英文摘要
Large Vision-Language Models (LVLMs) incur high computational costs from redundant visual tokens. Although training-free attention-based multi-layer pruning in the vision encoder stage has been explored as an effective strategy, we find that pruning in shallow layers consistently degrades performance. In this paper, we aim to understand this problem and seek a solution. By analyzing attention patterns across network depth, we find that shallow layers primarily function as edge detectors with chaotic attention maps, while deeper layers transition through local subject recognition and unstable semantic aggregation. To address the misalignment between pruning strategies and network stages, we propose STD, a hierarchical token pruning framework that adapts token selection mechanisms to the functional role of each network stage. STD employs High-Frequency Spectral Analysis in shallow layers to deterministically preserve structural edges, uses Gaussian-Smoothed Attention in intermediate layers to maintain spatial coherence, and introduces a Stability-Adaptive Trigger in deep layers to execute pruning only during semantically stable phases. Extensive experiments show that STD outperforms state-of-the-art pruning methods by 1.1% on LLaVA-1.5-7B with 88.9% token reduction, while also being plug-and-play and highly effective when combined with other methods, and by 2.1% on LLaVA-NeXT-7B with 94.4% reduction, delivering a 3.9x speed-up in the prefilling stage. Our code will be released at https://github.com/Twilight03/STD.
发表机构
- Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。