arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于更快、更鲁棒的视觉语言模型(VLM)的聚类与令牌去噪

Clustering and Token Denoising for Faster and More Robust VLMs

Baptiste Rossigneux, Inna Kucher, Vincent Lorrain, Emmanuel Casseau

arXiv 2608.19285首次发表:更新:

发表机构

Université Paris-Saclay; Univ. Rennes; Vsora(巴黎-萨克雷大学; 雷恩大学; 沃索拉公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出无需训练的ClustRS算法,通过聚类与残差收缩剪枝VLM的视觉令牌,在基准测试中实现显著性能提升,为高效抗噪VLM部署提供新方案。

AI 中文摘要

近期的视觉语言模型(VLM)通过在文本之外添加视觉令牌,增强了预训练大型语言模型(LLM)的能力,诸如LLaVA这类方法已展现出令人印象深刻的结果。然而,处理多达576或729个视觉令牌的计算负担,使得边缘部署颇具挑战性。尽管各类令牌剪枝技术需要重新训练,但也存在无需训练的方法,能轻松适配架构变更。我们提出ClustRS,这是一种用于鲁棒令牌剪枝的两部分、无需训练的算法。其第一部分是基于注意力权重的聚类算法,从每个语义簇中选择代表性令牌;第二部分是残差收缩,即对所选令牌进行的单次去噪步骤。这些无需训练的轻量步骤使LLaVA可适配真实世界数据,提升了对各类图像噪声类型与强度的鲁棒性。在ScienceQA-IMG和MM-VET基准上的实验结果表明,在LLaVA 1.5 7b模型上,我们的方法在极端噪声和令牌缩减条件下(令牌减少97%,降至16个令牌),比基于注意力和多样性的方法性能提升达20%;在LLaVA-OneVision上,在轻度噪声条件下,我们用不到基准三分之一的令牌即可达到与基准相当的性能。本研究证明了一种简单却强大的替代方案,可替代仅基于分数和仅基于多样性的剪枝规则,为计算高效且抗噪的VLM部署铺平了道路。

英文摘要

Recent Visual-Language Models (VLMs) have enhanced the capabilities of pre-trained LLMs by adding vision tokens alongside text, with approaches like LLaVA showing impressive results. However, the computational burden of processing up to 576 or 729 visual tokens makes edge deployment challenging. While various token pruning techniques require retraining, some are training-free and thus can easily adapt to architecture changes. We introduce ClustRS, a two-part, training-free algorithm for robust token pruning. Its first component is an attention-weighted, clustering algorithm that selects representative tokens from each semantic cluster. The second component, Residual Shrinkage, is a one-pass denoising step on the selected tokens. These training-free lightweight steps make LLaVA ready for real-world data, improving robustness to a wide range of image-noise types and intensities. Experimental results on the ScienceQA-IMG and MM-VET benchmarks show our method outperforms attention- and diversity-based methods by up to 20\% under extreme noise and token conditions (reducing tokens by 97\%, down to 16 tokens) on LLaVA 1.5 7b and achieves exceptional results on LLaVA-OneVision, where we match baseline performance with fewer than one-third of their tokens under mild noise conditions. Our study demonstrates a simple yet powerful alternative to both score-only and diversity-only pruning rules, paving the way for compute-efficient and noise-resilient VLM deployment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑