发表机构
Peking University; Nanyang Technological University, Singapore; University of Electronic Science and Technology of China; University of Michigan, Ann Arbor; De Artificial Intelligence Lab(北京大学; 南洋理工大学; 电子科技大学; 密歇根大学安娜堡分校; De人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
STAR-Pro提出免训练两阶段视觉令牌剪枝框架,通过自适应阶段保留广泛覆盖、渐进阶段按层注意力精炼,在LLaVA-Video-7B上减少90.5%令牌并保持92.7%性能。
AI 中文摘要
大型视觉-语言模型(LVLMs)实现了强大的多模态理解能力,但它们处理的数百至数千个视觉令牌带来了巨大的计算开销,这促使了免训练的视觉令牌剪枝方法的发展。在本工作中,我们对视觉令牌剪枝进行了两项互补的分析。首先,我们测量了跨模态融合前保留令牌的特征空间覆盖度,发现激进的剪枝会丢弃大量视觉信息。其次,我们追踪了解码器各层中文本到视觉的注意力,发现被视为重要的视觉令牌随深度变化显著,这使得一次性剪枝决策不可靠。综合这些发现,有效的剪枝应在融合前保留广泛的视觉覆盖,并在融合过程中随着跨模态证据的演变逐步精炼保留的令牌。因此,我们提出了STAR-Pro(阶段式自适应令牌缩减与渐进式精炼),一个免训练的两阶段框架。其自适应阶段应用枢轴QR分解来构建一个超预算的特征覆盖候选池,而其渐进阶段在选定的解码器层使用不断演变的文本到视觉注意力,在目标层平均令牌预算下剪枝出一个嵌套的幸存者集合。跨七个涵盖多种架构的LVLMs和18个图像及视频基准的广泛实验证明了STAR-Pro在激进剪枝下的有效性。在LLaVA-Video-7B上,STAR-Pro将视觉令牌减少了90.5%,保留了基线性能的92.7%,并实现了2.24倍的实测推理加速。代码可在该https URL获取。
英文摘要
Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose substantial computational overhead, motivating training-free visual token pruning. In this work, we conduct two complementary analyses of visual token pruning. First, we measure the feature-space coverage of tokens retained before cross-modal fusion and find that aggressive pruning discards substantial visual information. Second, we track text-to-visual attention across decoder layers and find that the visual tokens considered important change substantially with depth, making one-shot pruning decisions unreliable. Together, these findings show that effective pruning should preserve broad visual coverage before fusion and progressively refine the retained tokens as cross-modal evidence evolves during fusion. We therefore propose STAR-Pro (STage-Wise Adaptive Token Reduction with Progressive Refinement), a training-free two-stage framework. Its Adaptive Stage applies pivoted QR to construct an over-budget feature-coverage candidate pool, while its Progressive Stage uses evolving text-to-visual attention at selected decoder layers to prune a nested survivor set under a target layer-average token budget. Extensive experiments across seven LVLMs spanning multiple architectures and 18 image and video benchmarks demonstrate the effectiveness of STAR-Pro under aggressive pruning. On LLaVA-Video-7B, STAR-Pro reduces visual tokens by 90.5%, retains 92.7% of baseline performance, and achieves a $2.24\times$ measured inference speedup. Code is available at https://github.com/EasonAI-5589/starpro.
Comments26 pages, 7 figures