发表机构
National University of Singapore; InfRec, Cardinal AI Lab; The Hong Kong University of Science and Technology(新加坡国立大学; InfRec,Cardinal AI 实验室; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出VIP-Router,一种轻量级样本自适应路由方法,为每个输入动态选择最优视觉令牌剪枝策略,在VTC-Bench Group A上平均准确率相对提升26.9%,且即插即用、参数开销极低。
AI 中文摘要
多模态大语言模型(MLLMs)处理每张图像时需使用数百或数千个视觉令牌,导致高昂的推理成本。现有的视觉令牌剪枝方法虽能缓解这一开销,但它们隐含地假设单一固定的剪枝策略可统一应用于所有输入。我们的分析进一步揭示,按平均基准准确率对剪枝策略进行排序会掩盖显著的样本级互补性:尽管平均最优策略在整体上表现最佳,但替代策略在相当比例的个体样本上被证明更优。为利用这种多样性,我们提出VIP-Router,一种轻量级的视觉剪枝路由器,它能在指定剪枝水平下自适应地选择被预测为最适合每个输入的剪枝策略。基于低成本的视觉和文本特征,VIP-Router识别出最合适的候选策略,同时在预测剪枝不利时保留全令牌推理作为选项。在精心策划的剪枝敏感视觉感知基准套件VTC-Bench Group A上评估,VIP-Router在所有缩减比率下均持续优于最佳固定策略基线,平均准确率相对提升26.9%,在考虑实际令牌成本后平均效用相对提升22.0%。关键的是,VIP-Router以即插即用的方式运行,无需修改底层剪枝算法或模型权重,引入的可训练参数仅相当于主干的0.017%。此外,VIP-Router在各种MLLM主干上均证明有效,并在未见基准上产生一致增益,凸显了样本自适应路由在视觉令牌剪枝中的潜力。
英文摘要
Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, incurring prohibitive inference costs. While existing vision token pruning methods mitigate this overhead, they implicitly assume that a single fixed pruning strategy can be applied uniformly across all inputs. Our analysis further reveals that ranking pruning methods by average benchmark accuracy conceals substantial sample-wise complementarity: although the average-best strategy excels overall, alternative strategies prove superior on a significant fraction of individual samples. To harness this diversity, we propose VIP-Router, a lightweight VIsion Pruning Router that adaptively selects the pruning strategy predicted to be best suited to each input at a specified pruning level. Conditioned on low-cost visual and textual features, VIP-Router identifies the most suitable candidate strategy while retaining full-token inference as an option when pruning is predicted to be unfavorable. Evaluated on a curated suite of pruning-sensitive visual perception benchmarks, VTC-Bench Group A, VIP-Router consistently outperforms the best fixed strategy baseline across all reduction ratios, achieving a 26.9% relative improvement in average accuracy, and a 22.0% relative increase in average utility after accounting for realized token cost. Crucially, VIP-Router operates in a plug-and-play manner without modifying underlying pruning algorithms or model weights, introducing trainable parameters equivalent to merely 0.017\% of the backbone. Furthermore, VIP-Router proves effective across various MLLM backbones and yields consistent gains on unseen benchmarks, highlighting the potential of sample adaptive routing for visual token pruning.
Comments26 pages, 6 figures. Code will be released soon