arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26806cs.CV

多图像视觉语言模型中的多图像视觉令牌剪枝

Multi-Image Visual Token Pruning in Large Visual Language Models

Rongyang Zhang, Chengqiang Lu, Cong Li, Hongchao Gu, Tingjia Shen, Xuyang Zhi, Qimeng Wang, Yan Gao, Yi Wu, Yao Hu, Hao Wang, Enhong Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有LVLMs多图像剪枝方法的局限,提出无需训练的AVTP框架,在多类LVLMs上实现显著推理加速,同时保持较高准确率,部分模型性能甚至优于基线。

中文摘要 AI 辅助

随着现实应用中对多图像序列实时处理的需求不断增长,各类视觉令牌剪枝方法应运而生,以缓解大型视觉语言模型(LVLMs)面临的计算和上下文长度约束。然而,大多数现有剪枝方法依赖静态策略,难以适配不同架构的LVLMs和多图像场景,且额外受限于对注意力计算的依赖,无法与FlashAttention等高效技术兼容。为解决这些局限,本文提出一种无需训练的自适应视觉令牌剪枝(AVTP)框架,适用于各类LVLM架构。我们基于对不同LVLMs视觉注意力分布的实证分析,策略性确定剪枝层,并在多图像场景中实现自适应剪枝比例——重要性更高的图像会保留比例更多的令牌。我们在不同LVLMs上开展大量实验以验证AVTP的有效性和鲁棒性:具体而言,Qwen3VL-8B在多个多图像基准测试中实现2倍推理加速,同时保持96.1%的原始准确率;InternVL3.5-8B保留94.1%的准确率;LLaVA-OV-7B甚至超出其原始基线性能。我们的代码可通过此链接获取。

英文摘要

With the growing demand for processing multiple image sequences in real-world applications, various visual token pruning methods have emerged to mitigate the computational and context length constraints faced by Large Vision Language Models (LVLMs). However, most existing pruning approaches rely on static strategies that struggle to adapt across different architectural LVLMs and multi-image scenarios, and are additionally constrained by their dependence on attention computations that are incompatible with efficient techniques like FlashAttention. To address these limitations, we propose a training-free, Adaptive Visual Token Pruning (AVTP) framework, applicable to diverse LVLM architectures. We strategically determine pruning layers based on empirical analysis of visual attention distributions across various LVLMs, and implement adaptive pruning ratios in multi-image contexts where images of higher importance retain proportionally more tokens. We conduct extensive experiments across different LVLMs to demonstrate the effectiveness and robustness of AVTP. Specifically, Qwen3VL-8B achieves 2 times inference speedup while maintaining 96.1\% of its original accuracy on multiple multi-image benchmarks, InternVL3.5-8B retains 94.1\% accuracy, and LLaVA-OV-7B even exceeds its original baseline performance. Our code is available at \href{https://github.com/zry13/AVTP}{this link}.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Xiaohongshu Inc.(小红书科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑