arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于高分辨率多模态大语言模型中高效视觉令牌剪枝的结构化冗余建模

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs

Jouwon Song, Woohyeong Kim, Kyeongbo Kong

arXiv 2607.23046首次发表:更新:

发表机构

LG Electronics; Pusan National University(LG电子; 釜山国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高分辨率多模态大语言模型视觉令牌爆炸致延迟瓶颈问题,提出单向前向剪枝器SFPruner,通过语义引导岭杠杆方案和基于排名的方向掩码两个互补机制,在单次前向传递中实现冗余感知重要性选择,有效减少令牌选择时间并保持性能。

AI 中文摘要

近期的高分辨率多模态大语言模型每个输入会生成数千个视觉令牌,导致视觉令牌爆炸并带来严重的延迟瓶颈。虽然令牌剪枝可缓解此问题,但现有最优子集优化方法通常依赖迭代子集构建来兼顾视觉多样性和指令相关性。随着视觉令牌数量增加,这种顺序依赖性会带来显著的选择开销。为解决此局限,我们提出单向前向剪枝器(SFPruner),将冗余控制直接嵌入评分空间,通过两个互补机制在单次前向传递中实现冗余感知重要性选择。一是在协方差层面衰减冗余,引入语义引导的岭杠杆方案;二是基于排名的方向掩码通过不对称相似性竞争解决剩余重叠。大量评估表明,我们的方法保持稳定选择成本,在Qwen2.5-VL中,将512个令牌时的令牌选择过程从112.4毫秒减少到仅2.5毫秒,成功将理论令牌减少转化为实际推理加速,且在激进压缩下与现有技术相比保持极具竞争力的性能。

英文摘要

Recent high-resolution Multimodal Large Language Models (MLLMs) generate thousands of visual tokens per input, leading to a visual token explosion that introduces severe latency bottlenecks. While token pruning mitigates this issue, state-of-the-art subset-optimization methods typically rely on iterative subset construction to jointly capture visual diversity and instruction relevance. As visual token counts scale, this sequential dependency introduces significant selection overhead, severely limiting the translation of theoretical FLOPs reductions into actual wall-clock speedups. To address this limitation, we propose Single-Forward Pruner (SFPruner), a structural reformulation of visual token pruning that embeds redundancy control directly into the scoring space, bypassing the need for iterative combinatorial optimization. Our non-iterative framework achieves redundancy-aware importance selection in a single forward pass through two complementary mechanisms. First, to attenuate redundancy at the covariance level, we introduce a semantics-guided ridge leverage scheme. By integrating instruction relevance and visual saliency, this mechanism suppresses dominant covariance directions and mitigates representation bias. Second, ranking-based directional masking resolves residual overlap through asymmetric similarity competition, where higher-scoring tokens explicitly suppress redundant lower-scoring alternatives via parallel tensor operations. Extensive evaluations demonstrate that our approach maintains stable selection costs, reducing the token selection process by up to 110 ms, from 112.4 ms to just 2.5 ms at 512 tokens in Qwen2.5-VL. This structural efficiency successfully translates theoretical token reductions into tangible inference speedups while preserving highly competitive performance against state-of-the-art techniques under aggressive compression.

CommentsAccepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑