arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STAR-Pro:面向高效大型视觉-语言模型的阶段式令牌自适应缩减与渐进式精炼

STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models

Yichen Guo, Tinghao Wang, Qizhe Zhang, Lingbei Meng, Yuan Zhang, Jiajun Cao, Hao Jiang, Chenwei Wu, Jixian Wu, Sixiang Chen, Tao Luo, Hongyang Cheng, Kai Tang, Chenxi Li, Renyuan Li, Xiande Huang, Wenya Wang, Shanghang Zhang

arXiv 2609.05916首次发表:更新:

发表机构

Peking University; Nanyang Technological University, Singapore; University of Electronic Science and Technology of China; University of Michigan, Ann Arbor; De Artificial Intelligence Lab(北京大学; 南洋理工大学; 电子科技大学; 密歇根大学安娜堡分校; De人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

STAR-Pro提出免训练两阶段视觉令牌剪枝框架,通过自适应阶段保留广泛覆盖、渐进阶段按层注意力精炼,在LLaVA-Video-7B上减少90.5%令牌并保持92.7%性能。

AI 中文摘要

大型视觉-语言模型(LVLMs)实现了强大的多模态理解能力,但它们处理的数百至数千个视觉令牌带来了巨大的计算开销,这促使了免训练的视觉令牌剪枝方法的发展。在本工作中,我们对视觉令牌剪枝进行了两项互补的分析。首先,我们测量了跨模态融合前保留令牌的特征空间覆盖度,发现激进的剪枝会丢弃大量视觉信息。其次,我们追踪了解码器各层中文本到视觉的注意力,发现被视为重要的视觉令牌随深度变化显著,这使得一次性剪枝决策不可靠。综合这些发现,有效的剪枝应在融合前保留广泛的视觉覆盖,并在融合过程中随着跨模态证据的演变逐步精炼保留的令牌。因此,我们提出了STAR-Pro(阶段式自适应令牌缩减与渐进式精炼),一个免训练的两阶段框架。其自适应阶段应用枢轴QR分解来构建一个超预算的特征覆盖候选池,而其渐进阶段在选定的解码器层使用不断演变的文本到视觉注意力,在目标层平均令牌预算下剪枝出一个嵌套的幸存者集合。跨七个涵盖多种架构的LVLMs和18个图像及视频基准的广泛实验证明了STAR-Pro在激进剪枝下的有效性。在LLaVA-Video-7B上,STAR-Pro将视觉令牌减少了90.5%,保留了基线性能的92.7%,并实现了2.24倍的实测推理加速。代码可在该https URL获取。

英文摘要

Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose substantial computational overhead, motivating training-free visual token pruning. In this work, we conduct two complementary analyses of visual token pruning. First, we measure the feature-space coverage of tokens retained before cross-modal fusion and find that aggressive pruning discards substantial visual information. Second, we track text-to-visual attention across decoder layers and find that the visual tokens considered important change substantially with depth, making one-shot pruning decisions unreliable. Together, these findings show that effective pruning should preserve broad visual coverage before fusion and progressively refine the retained tokens as cross-modal evidence evolves during fusion. We therefore propose STAR-Pro (STage-Wise Adaptive Token Reduction with Progressive Refinement), a training-free two-stage framework. Its Adaptive Stage applies pivoted QR to construct an over-budget feature-coverage candidate pool, while its Progressive Stage uses evolving text-to-visual attention at selected decoder layers to prune a nested survivor set under a target layer-average token budget. Extensive experiments across seven LVLMs spanning multiple architectures and 18 image and video benchmarks demonstrate the effectiveness of STAR-Pro under aggressive pruning. On LLaVA-Video-7B, STAR-Pro reduces visual tokens by 90.5%, retains 92.7% of baseline performance, and achieves a $2.24\times$ measured inference speedup. Code is available at https://github.com/EasonAI-5589/starpro.

Comments26 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑