arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

注意尖峰:大型视觉-语言模型中视觉大规模激活的机制与脆弱性

Mind the Spike: Mechanisms and Brittleness of Visual Massive Activations in Large Vision-Language Models

Jonas Ngnawé, Yann Pequignot, Sabyasachi Sahoo, Christian Gagné, Frédéric Precioso, Sanmi Koyejo

arXiv 2609.32808首次发表:更新:

AI 中文总结

本研究揭示大型视觉-语言模型中视觉尖峰的机制与脆弱性,提出基于触发方向的攻击与预防干预,有效控制尖峰,覆盖25个模型。

AI 中文摘要

大型视觉-语言模型(LVLMs)从其纯文本基础模型中继承了大规模激活:尖峰现象,即少数固定的隐藏通道接收到的值比典型幅度高出数千倍。文本尖峰系统性地出现在早期层的固定初始位置,与输入内容无关。视觉尖峰因图像而异,但其形成是否在LVLMs中遵循一致的模式,以及它们对图像扰动的响应方式,仍是未解问题。我们发现,一些LVLMs不形成视觉尖峰,而其他模型则以不同的速率形成尖峰,通常出现在较深的层。我们从模型权重中识别出触发方向,并得出一个可解释的位置规则:在语言模型解码器运行之前,最终的尖峰标记主要局限于那些与图像其余部分共享最少的标记。至关重要的是,视觉尖峰极其脆弱。常见的损坏经常产生和重新定位尖峰,较少移除它们,从而提高总体发生率。我们基于触发引导的尖峰攻击在小的ℓ∞预算下故意创建或移除尖峰,在十个产生尖峰的模型中,九个仅需1/255就足够。最后,我们的预防性干预在尖峰爆发前仅移除触发组件,在干净和受扰动的图像上消除或大幅减少尖峰,同时几乎不改变其他图像标记。我们的研究涵盖了基于10个家族的18个已发布纯文本基础模型构建的25个基于适配器的LVLMs,参数范围从2B到72B。

英文摘要

Large vision-language models (LVLMs) inherit massive activations from their text-only bases: spikes where a few fixed hidden channels receive values thousands of times above the typical magnitude. The text spike systematically appears in early layers at a fixed initial position, independently of input content. Visual spikes vary across images, but whether their formation follows a consistent pattern across LVLMs and how they respond to image perturbations remain open questions. We find that some LVLMs do not form visual spikes, while others spike at different rates, typically in deeper layers. We identify the trigger direction from model weights and an interpretable location rule: before the language model decoder runs, eventual spike tokens are largely restricted to those sharing least with the rest of the image. Crucially, visual spikes are strikingly brittle. Common corruptions frequently create and relocate spikes, and less often remove them, raising overall incidence. Our trigger-guided spike attack deliberately creates or removes spikes under a small $\ell_\infty$ budget, with 1/255 enough in nine of the ten models that spike. Finally, our preventive intervention removes only the trigger component before spikes erupt, eliminating or substantially reducing spikes on clean and perturbed images while leaving the other image tokens nearly unchanged. Our study spans 25 adapter-based LVLMs built on 18 released text-only bases from 10 families, ranging from 2B to 72B parameters.

Comments58 pages, 15 figures, 43 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑