arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19990cs.CV

QCPruner:面向视觉令牌剪枝的查询条件化群体覆盖

QCPruner: Query-Conditioned Population Coverage for Visual Token Pruning

  • Guizhou University(贵州大学)
  • Shanghai Ocean University(上海海洋大学)

机构由 AI 辅助整理,请以论文原文为准。

Shengli He, Yongchao Liang, Roumeng He, Junjie Zeng, Jiyuan He, Can Wu, Li Zheng

AI总结:

QCPruner通过查询条件化双边效用加权,实现无需训练的视觉令牌剪枝,在多个MLLM上以子模贪心保证取得最高平均相对性能。

AI中文摘要:

多模态大语言模型(MLLMs)中高视觉令牌负载促使无需训练的剪枝方法以减少后期层计算,但在固定预算下,剪枝必须保留与查询相关的证据,同时避免冗余。现有方法对令牌进行排序、多样化所选子集或优化覆盖,但未使用共享的逐视觉查询效用对视觉目标和候选代表进行加权。我们提出QCPruner,通过双边效用加权使两个角色均查询条件化。利用关键词匹配的查询锚点,QCPruner将两种跨模态线索融合为效用,并将其应用于视觉亲和力覆盖中的视觉目标和候选代表。由此产生的非负设施位置目标函数是单调且子模的,保留标准(1-1/e)贪心保证,且无需模型训练或参数更新。在LLaVA-1.5、LLaVA-NeXT、LLaVA-Video和Qwen2.5-VL上,QCPruner在报告的每个令牌预算下均达到所评估完整系统剪枝方法中的最高平均相对性能。在LLaVA-1.5-7B的576个令牌中剪枝至32个时,它保留了未剪枝性能的96.1%,而最强评估基线为93.9%。在Qwen2.5-VL-7B的1296个令牌中剪枝至256个时,相应数值为96.7%和92.5%。

英文摘要:

The high visual-token load in multimodal large language models (MLLMs) motivates training-free pruning to reduce later-layer computation, but under a fixed budget, pruning must preserve query-relevant evidence while avoiding redundancy. Existing methods rank tokens, diversify selected subsets, or optimize coverage without using a shared per-visual query utility to weight both visual targets and candidate representatives. We introduce QCPruner, which makes both roles query-conditioned through bilateral utility weighting. Using keyword-matched query anchors, QCPruner fuses two cross-modal cues into utility and applies it to both visual targets and candidate representatives within visual-affinity-based coverage. The resulting nonnegative facility-location objective is monotone and submodular, retains the standard (1-1/e) greedy guarantee, and requires no model training or parameter updates. Across LLaVA-1.5, LLaVA-NeXT, LLaVA-Video, and Qwen2.5-VL, QCPruner achieves the highest average relative performance among evaluated complete-system pruning methods at every reported token budget. At 32 of 576 tokens on LLaVA-1.5-7B, it retains 96.1% of unpruned performance, versus 93.9% for the strongest evaluated baseline. At 256 of 1296 tokens on Qwen2.5-VL-7B, the corresponding values are 96.7% and 92.5%.

↑