arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

V-CoLA:基于线性注意力的视觉令牌压缩

V-CoLA: Vision Token Compression with Linear Attention

Hao Jiang, Yiru Mao, Tianpeng Bu, Hao Zhou, Hongtao Duan, Wang Jing, Bowen Xu, Xin Chen, Lulu Hu, Bin Yang, Yongliang Tao, Minying Zhang

arXiv 2610.11251首次发表:更新:

发表机构

Alibaba Cloud Computing, Alibaba Group(阿里巴巴集团阿里云计算)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉-语言模型因视觉令牌主导输入序列导致计算开销大的问题,提出专为线性注意力设计的无训练压缩框架V-CoLA,通过感知唯一性的重要性准则与自适应令牌合并策略,在大幅减少视觉令牌的同时保持高性能并提升预填充速度。

AI 中文摘要

视觉-语言模型(VLMs)展现出强大能力,但存在显著计算开销,因为视觉令牌主导输入序列,这促使视觉令牌压缩成为缓解该负担的关键方向。然而,随着融入线性注意力的混合架构(如Qwen3.5)出现,专为softmax注意力设计的现有方法难以泛化。我们的分析表明,基于注意力和相似度的方法均出现明显性能下降,凸显了针对该场景定制压缩方法的迫切需求。为此,我们提出V-CoLA,一种专为线性注意力设计的高效无训练令牌压缩框架。V-CoLA引入了一种新颖的“感知唯一性的重要性准则”以识别关键视觉令牌,结合执行压缩的“自适应令牌合并策略”。所有组件在实现层面优化,以保持与线性注意力的分块并行性兼容,确保强实用价值。在多个基准上的大量实验证明了V-CoLA的优越性:仅使用50.0%的视觉令牌即可达到原始性能的99.5%,仅使用12.5%时仍能达到超过88.0%的性能,同时实现1.86倍至6.15倍的预填充加速。

英文摘要

Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence. This motivates vision token compression as a key direction to alleviate the burden. However, with the emergence of hybrid architectures incorporating linear attention (\eg, Qwen3.5), prior methods designed for softmax attention struggle to generalize. Our analysis reveals that both attention- and similarity-based approaches suffer notable performance degradation, underscoring the urgent need for compression methods tailored to this regime. To this end, we propose \textbf{V-CoLA}, an efficient training-free token compression framework specifically designed for linear attention. V-CoLA introduces a novel \textit{uniqueness-aware importance criterion} for identifying critical vision tokens, coupled with an \textit{adaptive token merging strategy} that performs compression. All components are optimized at the implementation level to remain compatible with the chunk-wise parallelism of linear attention, ensuring strong practical value. Extensive experiments across multiple benchmarks demonstrate the superiority of V-CoLA: it achieves 99.5\% of the original performance with only 50.0\% of vision tokens, and over 88.0\% with as few as 12.5\%, while delivering a 1.86$\times$ to 6.15$\times$ prefill speedup.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑