arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

消息而非令牌:用于忠实VLM压缩的基础核心集

Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression

Long Qian, Jiaqi Wei, Bingke Zhu, Yingying Chen, Jinqiao Wang

arXiv 2608.02134首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Wuhan AI Research; Cardiff University(中国科学院自动化研究所; 中国科学院大学; 武汉人工智能研究院; 卡迪夫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对VLM视觉令牌压缩的现有缺陷,提出无需训练的GMC方法,通过联合分配多维度证据支持并传输被丢弃状态,在大幅减少视觉令牌的同时保持VLM性能,经多基准实验验证有效。

AI 中文摘要

现代视觉语言模型(VLMs)将高分辨率图像转换为长序列的视觉令牌,每个令牌都会经过语言解码器并保留在其提示KV缓存中,这增加了推理成本,促使人们采取激进的视觉压缩方法。现有的基于分数的方法为每个令牌分配独立的重要性分数并保留前K个令牌,但文本查询消耗的是视觉群体的集体、带符号的注意力消息,而非孤立的图像块,因此同等大小的前K集合可能会反复覆盖一个显著区域、遗漏稀疏但互补的证据,并丢弃被移除群体所携带的信息。为此,我们将忠实的视觉压缩表述为为解码器消息构建紧凑核心集的问题,并引入无需训练的基础消息核心集剪枝(GMC)方法,该方法会在查询基础、外观和坐标感知证据之间联合分配支持,随后在物理压缩和原生注意力恢复前,将被丢弃的状态传输到选定代表的原始多模态位置。这将忠实压缩分解为两个耦合组件,包括选择覆盖所需消息模式的载体,以及在这些载体上实现带符号的群体消息。我们进一步推导了将其误差与带符号消息失真、视觉创新性和候选边际稳定性关联起来的边界。在多个VLM系列和不同基准上的实验显示出强大性能:GMC-H2在Qwen2.5-VL-7B上保留了97.78%的全相对平均能力,同时减少了80.2%的视觉令牌;GMC-L16则达到了100.36%。受控干预验证了集体支持和群体实现共同推动了这些增益。

英文摘要

Modern vision language models (VLMs) turn high-resolution images into long sequences of visual tokens. Every token traverses the language decoder and persists in its prompt KV cache, inflating inference cost and motivating aggressive visual compression. Existing score-based methods assign each token an independent importance score and retain the Top-K. However, text queries consume collective, signed attention messages from the visual population, not isolated patches. Consequently, equally sized Top-K sets can repeatedly cover one salient region, omit sparse but complementary evidence and discard information carried by the removed population. We therefore formulate faithful visual compression as constructing a compact coreset for decoder messages, and introduce our training-free Grounded Message Coreset Pruning (GMC) which jointly allocates support across query-grounded, appearance, and coordinate-aware evidence, then transports discarded states into selected representatives at their original multimodal positions before physical compaction and native attention resume. This decomposes faithful compression into two coupled components, including selecting carriers that cover the required message modes and realizing the signed population message on those carriers. We further derive bounds connecting their errors to signed-message distortion, visual innovation, and candidate-margin stability. Experiments across multiple VLM families and diverse benchmarks demonstrate strong performance, with GMC-H2 retaining 97.78% Full-relative mean capability on Qwen2.5-VL-7B using 80.2% fewer visual tokens, while GMC-L16 reaches 100.36%. Controlled interventions verify that collective support and population realization jointly drive these gains.

Comments32 pages, 6 figures, 18 tables, including appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑