发表机构
Digital Environment Research Institute (DERI), Queen Mary University of London; Shanghai Institute of Artificial Intelligence for Education, East China Normal University; National Institute of Education, Nanyang Technological University; School of Computer Science and Engineering, Jiangsu University of Science and Technology(伦敦大学玛丽女王学院数字环境研究所(DERI); 华东师范大学上海人工智能教育研究院; 南洋理工大学国家教育学院; 江苏科技大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出无需额外参数或微调的E2S-Pruner,通过两阶段证据融合实现视觉-语言模型的视觉令牌剪枝,在LLaVA-1.5-7B上显著提升吞吐量并保留性能,且具跨模型泛化性。
AI 中文摘要
视觉-语言模型通常将图像编码为数百个视觉令牌,导致推理延迟和GPU内存开销显著。现有剪枝方法大多依赖注意力分数,直接聚合注意力头和网络层的输出,难以表征证据不确定性和冲突。本文提出E2S-Pruner,一种无需辅助模型、可训练参数或微调的视觉令牌剪枝渐进式两阶段证据融合框架。第一阶段,E2S-Pruner将每个注意力头视为独立证据源,从证据清晰度和头间一致性估计其可靠性,并用重要、不重要、不确定三种状态表示每个视觉令牌。第二阶段,采用Dempster–Shafer证据理论量化层间冲突并融合多个网络层的互补证据。本文进一步引入空间新颖性约束,以促进覆盖不同图像区域,防止保留的令牌集中在少数局部显著区域。在LLaVA-1.5-7B上,当平均保留视觉令牌数为192、128、64时,E2S-Pruner分别保留98.0%、96.8%、90.6%的聚合性能,同时在128令牌和64令牌设置下,吞吐量分别提升1.96倍和2.09倍。在Qwen2-VL-7B上的实验进一步证明了跨模型泛化能力。代码可在该URL获取。
英文摘要
Vision-language models typically encode an image into hundreds of visual tokens, incurring substantial inference latency and GPU memory overhead. Existing pruning methods largely rely on attention scores and directly aggregate outputs across attention heads and network layers, making it difficult to characterize evidential uncertainty and conflict. We propose E2S-Pruner, a progressive two-stage evidence-fusion framework for visual token pruning that requires no auxiliary model, trainable parameters, or fine-tuning. In the first stage, E2S-Pruner treats each attention head as an independent evidence source, estimates its reliability from evidence clarity and inter-head consistency, and represents each visual token using three states: important, unimportant, and uncertain. In the second stage, Dempster--Shafer evidence theory is used to quantify inter-layer conflict and fuse complementary evidence from multiple network layers. We further introduce a spatial novelty constraint that promotes coverage of distinct image regions and prevents the retained tokens from concentrating in a few locally salient areas. On LLaVA-1.5-7B, E2S-Pruner retains 98.0%, 96.8%, and 90.6% of the aggregate performance when the average numbers of retained visual tokens are 192, 128, and 64, respectively, while improving throughput by 1.96x and 2.09x under the 128-token and 64-token settings. Experiments on Qwen2-VL-7B further demonstrate cross-model generalization. Code is available at https://github.com/taoyu-qian/E2S-Pruner.git.