视觉语言模型(VLMs)是否在不同模态间共享安全神经元?
Do VLMs Share Safety Neurons Across Modalities?
- SB Intuitions Corp.(SB Intuitions公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对10个视觉语言模型,通过因果神经元层面分析,提出两阶段迭代消融检测流程与两个模态隔离基准,发现文本安全神经元集中、视觉安全高维弥散的差距,解释了当前对齐未缩小视觉安全差距的原因。
AI中文摘要:
视觉语言模型(VLMs)能够顺从通过图像传递的有害请求,即便其大型语言模型(LLM)主干会拒绝相同的文本内容。尽管先前的研究从经验或表征层面描述了这些越狱现象,但视觉输入如何在神经元层面干扰安全通路仍未被探索。我们通过对10个VLMs中安全机制的因果神经元层面分析填补了这一空白。我们提出了一种考虑自我修复的迭代消融两阶段检测流程,并引入了两个分离视觉与文本安全信号的模态隔离基准:ViSafe-Detect和ViSafe-Eval。我们的分析揭示:(i)VLMs中的文本安全是可定位的:约88个神经元(占比<0.01%),对其进行针对性消融会大幅降低拒绝行为;(ii)文本安全神经元构成了主要的拒绝通路:消融这些神经元是唯一在所有模型中都能一致且大幅降低拒绝行为的干预措施;(iii)视觉安全在单神经元层面是高维且弥散的:文本安全集中在约5个子空间方向,而视觉安全则需要≥50个。这一差距在不同架构中均存在,解释了当前对齐为何未能缩小视觉安全差距。项目页面位于:this https URL 警告:本文可能包含有害内容示例。
英文摘要:
Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content in text. While prior work characterizes these jailbreaks empirically or at the representation level, how visual inputs perturb safety pathways at the neuron level remains uncharted. We close this gap with a causal, neuron-level analysis of safety mechanisms in 10 VLMs. We propose a two-stage detection pipeline with iterative ablation that accounts for self-repair, and introduce two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval, which decouple visual and textual safety signals. Our analysis reveals: (i) Text safety in VLMs is localizable: $\sim$88 neurons ($<$0.01%) whose targeted ablation substantially reduces refusal. (ii) Text safety neurons constitute the dominant refusal pathway: ablating them is the only intervention that consistently and substantially reduces refusal across all models. (iii) Visual safety is high-dimensional and diffuse at the single-neuron level: text safety concentrates in $\sim$5 subspace directions while visual safety requires $\geq$50. This gap holds across architectures, explaining why current alignment has not closed the visual safety gap. Project page is at: https://jiaxuan-li.github.io/vlm-safety-neuron/ Warning: this paper may include examples of harmful content.