发表机构
The Hong Kong University of Science and Technology (Guangzhou); Tencent Yuanbao; ARC Lab, Tencent(香港科技大学(广州); 腾讯元宝; 腾讯ARC实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLM重排序器的令牌剪枝偏差问题,提出无需训练的RaDiCal框架,在多检索基准和VLM架构上实现高效剪枝,降低计算量并提升速度,同时保持排序性能。
AI 中文摘要
用作列表式重排序器的大型视觉语言模型(VLM)必须联合处理每个查询的数十个候选对象的视觉令牌,因此令牌剪枝对于实际部署至关重要。现有的剪枝方法通过注意力显著性保留令牌,但我们发现显著性与排序贡献存在系统性偏差:视觉上突出的令牌通常捕获候选对象间共有的与顺序无关的模式。这种偏差具有层依赖性:仅在注意力集中的区域,显著性才具有信息价值,而归一化注意力熵可诊断这种可靠性偏移(皮尔逊相关系数r=0.87)。我们提出RaDiCal(秩判别校准,Rank-Discriminative Calibration),这是一种无需训练的框架,它利用归一化注意力熵确定何时可信任显著性,将其与无注意力的秩判别先验融合,并从同一信任范围内选择剪枝层。在三个检索基准和多个VLM架构上,RaDiCal在令牌预算为20%时,Flickr30K上的MRR@10与Dense相当,在MSCOCO上超越Dense;在FashionIQ上的所有剪枝方法中排名第一;在保留率为10%时,Flickr30K和MSCOCO上的结果与Dense的差距在1.2个百分点以内。它将FLOPs降低39%-45%,在两种VLM架构上实现1.28-1.45倍的实测加速,且无需针对特定数据集重新调整。
英文摘要
Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning essential for practical deployment. Existing pruning methods retain tokens by attention saliency, yet we show that saliency is systematically misaligned with ranking contribution: visually prominent tokens often capture order-neutral patterns shared across candidates. This mismatch is layer-dependent: saliency becomes informative only where attention is concentrated, and normalized attention entropy diagnoses the reliability shift (Pearson r=0.87). We propose RaDiCal (Rank-Discriminative Calibration), a training-free framework that uses normalized attention entropy to decide when saliency can be trusted, fusing it with an attention-free rank-discriminative prior and selecting pruning layers from the same trust landscape. Across three retrieval benchmarks and multiple VLM architectures, RaDiCal matches Dense MRR@10 on Flickr30K and surpasses it on MSCOCO at a 20% token budget, ranks first among all pruning methods on FashionIQ, and holds within 1.2 pp on Flickr30K and MSCOCO at 10% retention. It cuts FLOPs by 39--45% and delivers 1.28--1.45$\times$ measured speedups across two VLM architectures without dataset-specific retuning.
CommentsEMNLP 2026 (Main Conference)