arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

层感知位置嵌入用于多模态大语言模型中的视觉令牌剪枝

Layer-Aware Position Embeddings for Visual Token Pruning in Multimodal Large Language Models

Yahong Wang, Zhangkai Ni, Juncheng Wu, Yuyin Zhou, Ying Wen, Lianghua He

arXiv 2609.23715首次发表:更新:

发表机构

Tongji University; University of California, Santa Cruz; East China Normal University; Shanghai Eye Disease Prevention and Treatment Center(同济大学; 加州大学圣克鲁兹分校; 华东师范大学; 上海市眼病防治中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型视觉令牌剪枝中位置嵌入的局限,提出层感知位置嵌入策略,在定位敏感层用稀疏嵌入、其余层用连续嵌入,提升剪枝模型的综合多模态性能。

AI 中文摘要

多模态大语言模型(MLLMs)由于依赖数百个视觉令牌来表示图像,产生了大量的计算开销。虽然令牌剪枝已成为降低MLLMs推理成本的一种有前景的方法,但现有方法通常使用稀疏或连续的位置嵌入来重新分配保留令牌的位置,每种方法都引入了不同的局限性。稀疏位置嵌入往往会降低分配给视觉令牌的注意力值,从而削弱MLLMs的感知能力,而连续位置嵌入则会破坏视觉令牌的原始空间对应关系,导致定位能力减弱。为了缓解这一问题,我们对语言解码器进行了逐层分析,并观察到中间层在令牌剪枝下对维持MLLMs的定位能力起着关键作用。基于这一观察,我们提出了一种层感知位置嵌入策略,该策略在定位敏感层切换到稀疏位置嵌入,而在其他层保持连续位置嵌入。在代表性剪枝方法和多样化基准上的大量实验表明,与标准的稀疏和连续位置嵌入相比,我们的方法提高了剪枝后MLLMs的综合多模态性能。

英文摘要

Multimodal large language models (MLLMs) incur substantial computational overhead due to the reliance on hundreds of visual tokens to represent images. While token pruning has emerged as a promising approach to reduce the inference cost of MLLMs, existing methods typically reassign position embeddings to the retained tokens using either sparse or continuous position embeddings, each introducing distinct limitations. Sparse position embeddings tend to decrease the attention value allocated to visual tokens, thereby degrading the perception capability of MLLMs, whereas continuous position embeddings disrupt the original spatial correspondence of visual tokens, leading to weakened grounding capability. To mitigate this issue, we perform layer-wise analysis of the language decoder and observe that intermediate layers play a critical role for maintaining the grounding capability of MLLMs under token pruning. Based on this observation, we propose a layer-aware position embedding strategy, which switches to sparse position embeddings at grounding-sensitive layers while maintaining continuous position embeddings elsewhere. Extensive experiments across representative pruning methods and diverse benchmarks demonstrate that our approach improves the comprehensive multimodal performance of pruned MLLMs compared with standard sparse and continuous position embeddings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑