StepPrune:多模态大语言模型中的自适应序列视觉标记选择
StepPrune: Adaptive Sequential Visual Token Selection across Multimodal Large Language Models
浏览论文内容
中文总结 AI 辅助
StepPrune通过自适应序列决策和STOP动作动态选择视觉标记,在多种MLLM上实现高效剪枝,保留性能的同时显著加速推理。
中文摘要 AI 辅助
在多模态大语言模型(MLLMs)中,视觉前缀占每层计算量的主要部分,因此视觉标记剪枝成为加速推理的直接方法。现有的top-K方法通常独立评估标记,并对所有输入应用统一的预算,忽略了选择之间的依赖关系以及样本间视觉复杂度的差异。相比之下,我们提出StepPrune,将视觉标记剪枝建模为自适应序列决策过程。基于先前选择的标记和文本上下文,StepPrune逐步构建保留子集,并通过学习到的STOP动作自动确定其大小。在训练期间,方差保持噪声门为离散选择过程提供可微的替代,而在推理期间,未选择的标记在语言模型预填充之前被物理移除。分组选择机制进一步将StepPrune扩展到高分辨率输入。在LLaVA-1.5、LLaVA-NeXT、Qwen2.5-VL和InternVL3上的实验表明,StepPrune在LLaVA-1.5、Qwen2.5-VL和InternVL3上所有评估剪枝率下均实现了最佳的平均归一化性能保留,同时在LLaVA-NeXT的显著更长的AnyRes前缀上保持竞争力。在LLaVA-1.5上,StepPrune在剪除88.9%的视觉标记时保留了完整前缀归一化性能的94.6%。在平均保留数量为64时,StepPrune将预填充延迟从59.95毫秒降低到40.05毫秒,对应1.50倍的预填充加速。
英文摘要
Visual prefixes account for a major portion of the per-layer computation in multimodal large language models (MLLMs), making visual-token pruning a direct approach to accelerating inference. Existing top-K methods typically evaluate tokens independently and apply a uniform budget to all inputs, overlooking both selection-dependent interactions and variations in visual complexity across samples. In contrast, we propose StepPrune, which formulates visual-token pruning as an adaptive sequential decision process. Conditioned on previously selected tokens and textual context, StepPrune progressively constructs the retained subset and automatically determines its size through a learned STOP action. During training, a variance-preserving noise gate provides a differentiable surrogate for the discrete selection process, whereas during inference, unselected tokens are physically removed before language-model prefill. A grouped selection mechanism further extends StepPrune to high-resolution inputs. Experiments across LLaVA-1.5, LLaVA-NeXT, Qwen2.5-VL, and InternVL3 show that StepPrune achieves the best average normalized performance retention across all evaluated pruning rates on LLaVA-1.5, Qwen2.5-VL, and InternVL3, while remaining competitive on the substantially longer AnyRes prefixes of LLaVA-NeXT. On LLaVA-1.5, StepPrune retains 94.6% of the full-prefix normalized performance while pruning 88.9% of the visual tokens. At a mean retained count of 64, StepPrune reduces prefill latency from 59.95 ms to 40.05 ms, corresponding to a 1.50x prefill speed-up.
发表机构
- Shenzhen University of Advanced Technology(深圳先进技术大学)
- Nanjing University(南京大学)
- CUHK MMLab, CPII under InnoHK(香港中文大学多媒体实验室,香港物流及供应链管理应用技术研发中心(InnoHK))
机构由 AI 辅助整理,请以论文原文为准。