VPRune:高效的无训练大语言模型前视觉令牌剪枝
VPRune: Efficient Training-free Pre-LLM Visual Token Pruning
浏览论文内容
中文总结 AI 辅助
VPRune提出无训练预LLM视觉令牌剪枝框架,通过多样性选择、令牌回收和位置恢复,在FastVLM-1.5B上实现精度与压缩的良好平衡,并降低边缘设备推理延迟。
中文摘要 AI 辅助
视觉令牌剪枝是降低大型视觉-语言模型(LVLMs)推理成本的一种有前景的方法,然而激进的令牌缩减往往会导致显著的性能下降。我们识别出导致这种下降的三个关键因素:文本引导的选择偏差、丢弃令牌造成的信息损失,以及序列压缩引起的位置失真。基于这些观察,我们提出了VPRune,一个无训练的、在大语言模型之前的剪枝框架,由纯视觉多样性选择、相似性引导的令牌回收和位置保持恢复组成。在FastVLM-1.5B上跨多个视觉-语言基准的实验表明,VPRune实现了有利的精度-压缩权衡,在激进压缩下优势尤为明显。此外,在边缘设备上的评估显示,VPRune有效降低了端到端推理延迟,同时保持了优越的任务性能,展示了其在资源受限的LVLM部署中的实用性。
英文摘要
Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We identify three key factors behind this degradation: text-guided selection bias, information loss from discarded tokens, and positional distortion caused by sequence compaction. Based on these observations, we propose \textbf{VPRune}, a training-free pre-LLM pruning framework consisting of visual-only diversity selection, similarity-guided token recycling, and position-preserving restoration. Experiments on FastVLM-1.5B across multiple vision-language benchmarks demonstrate that VPRune achieves a favorable accuracy--compression trade-off, with particularly pronounced advantages under aggressive compression. Furthermore, evaluations on edge-device show that VPRune effectively reduces end-to-end inference latency while maintaining superior task performance, demonstrating its practicality for resource-constrained LVLM deployment.
发表机构
- Northeastern University(东北大学)
- Beihang University(北京航空航天大学)
- Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。