arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VisCo:利用大语言模型作为视觉令牌压缩的内在编码器

VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression

Yupeng Zheng, Kai Zou, Bin Liu, Nenghai Yu

arXiv 2607.12756首次发表:更新:

AI 中文总结

研究针对视觉语言模型处理视觉令牌时的高成本问题,提出训练高效的自压缩框架VisCo,将预训练VLM用作内在压缩器,通过参数共享自动编码器压缩视觉信息,实验证明其在压缩率上超先前方法,还能捕获互补表示改进基础模型。

AI 中文摘要

视觉语言模型(VLM)处理大量视觉令牌,导致推理延迟和内存开销巨大,促使人们对视觉令牌压缩展开广泛研究。无训练策略依赖启发式指标,在高压缩率下性能大幅下降;许多基于训练的方法引入外部压缩模块,导致大量重新训练成本并损害VLM的先验知识。有效的视觉令牌压缩依赖于强大的信息编码,而预训练的VLM已具备此能力但未被现有方法充分利用。基于此,我们提出了VisCo,这是一个训练高效的自压缩框架,将预训练的VLM本身重新用作内在压缩器。VisCo是一个参数共享自动编码器,使用少量内存令牌压缩视觉信息,并将分层信息从编码传输到解码。实验表明,VisCo在所有评估的压缩率上都超过了先前的方法,在更激进的压缩下有更大的提升,甚至在极端的单令牌设置下也保持稳定。此外,当与原始视觉令牌结合时,学习到的内存令牌甚至可以改进基础模型,这表明VisCo除了压缩之外还捕获了互补表示。

英文摘要

Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated extensive research on visual token compression. While training-free strategies rely on heuristic metrics and suffer significant performance degradation under high compression ratios, many training-based methods introduce external compression modules that force the VLM backbone to adapt, incurring substantial retraining cost and compromising VLMs' priors. Effective visual token compression hinges on strong information encoding, a capability already present in pretrained VLMs but underutilized by existing approaches. Motivated by this, we propose VisCo, a training-efficient self-compression framework that reuses the pretrained VLM itself as an intrinsic compressor. VisCo is a parameter-sharing autoencoder that compresses visual information using a small set of memory tokens and transfers hierarchical information from encoding to decoding. Experiments show that VisCo surpasses prior methods across all evaluated compression ratios, with larger gains under more aggressive compression, and remains stable even in the extreme single-token setting. Moreover, when combined with the original visual tokens, the learned memory tokens can even improve the base model, suggesting that VisCo captures complementary representations beyond compression. Code is available at: \href{https://github.com/Zyvpeng/VisCo}{\textcolor{blue}

CommentsAccepted by ACM MM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑