arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33397cs.AI

CoViST:通过可组合状态进行视觉令牌压缩

CoViST: Visual Token Compression via Composable States

Qi Zhang, Xiandong Meng, Ronggang Wang, Siwei Ma

首次发表
浏览论文内容

中文总结 AI 辅助

CoViST通过可组合视觉状态实现免训练的视觉令牌压缩,保留位置和贡献信息,在七个LLaVA基准上以低预算保持高性能。

中文摘要 AI 辅助

视觉令牌压缩通过用更少的令牌表示图像来降低视觉语言模型的推理成本。然而,大多数现有方法将视觉令牌压缩为减少的集合,使得每个令牌所代表的视觉证据量及其原始空间上下文保持隐式。因此,压缩表示并未明确编码每个代表携带多少视觉信息或其在原始图像中的位置。这种限制即使在单次缩减后也会出现,并且当压缩在解码器层中重复进行时变得更加明显。为解决此问题,我们提出CoViST,一种免训练框架,将压缩图像表示为可组合的视觉状态。具体而言,该状态将代表性特征与原始位置、有效贡献权重和可重用选择元数据相结合。CoViST通过覆盖引导选择和基于守恒的贡献组合来构建此状态,并明确将其贡献和位置信息纳入解码器注意力中。状态的每个组件在连续缩减下保持其解释,使得相同的公式能够支持预填充前的固定压缩和解码器内的渐进压缩。在七个LLaVA-1.5-7B基准上的实验结果表明,CoViST-Fixed在192、128和64个令牌下分别保留了未压缩性能的99.9%、99.5%和98.1%,而CoViST-Pro在相应的层平均预算下分别保留了99.8%、99.9%和99.1%,在各自预算设置下优于最先进的方法。代码将公开发布。

英文摘要

Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. Therefore, the compressed representation does not explicitly encode how much visual information each representative carries or where it lies in the original image. This limitation arises even after a single reduction and becomes more pronounced when compression is repeated across decoder layers. To address this issue, we propose CoViST, a training-free framework that represents a compressed image as a composable visual state. Specifically, the state combines representative features with original positions, effective contribution weights, and reusable selection metadata. CoViST constructs this state through coverage-guided selection and conservation-based contribution composition, and explicitly incorporates its contribution and positional information into decoder attention. Each component of the state retains its interpretation under successive reductions, enabling the same formulation to support both fixed compression before prefill and progressive compression within the decoder. Experimental results on seven LLaVA-1.5-7B benchmarks show that CoViST-Fixed retains 99.9\%, 99.5\%, and 98.1\% of uncompressed performance at 192, 128, and 64 tokens, respectively, and CoViST-Pro retains 99.8\%, 99.9\%, and 99.1\% at the corresponding layer-average budgets, outperforming state-of-the-art methods under their respective budget settings. Code will be released publicly.

发表机构

  • Peng Cheng Laboratory(鹏城实验室)
  • Peking University Shenzhen Graduate School(北京大学深圳研究生院)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑