AI 中文总结
SPARK通过在预填充阶段选择性修复KV记忆中的危害相关子空间,无需修改参数或推理时分类器,显著降低多模态越狱攻击成功率,同时保持模型通用能力。
AI 中文摘要
视觉语言模型(VLM)仍然容易受到越狱攻击,这些攻击将有害意图分布在文本和图像中,使得单模态安全机制不足。我们研究是否可以在预填充阶段形成的多模态键值(KV)记忆中直接缓解这种漏洞,而无需在推理时修改模型参数。我们提出了SPARK,一个用于针对性KV记忆修复的两阶段框架。第一阶段使用一次性诊断适配器来识别多模态键和值表示中与危害相关的方向。第二阶段将这些方向投影出去,学习一个轻量级的残差修复,并用图像结构先验锚定修复后的键以保持视觉接地。SPARK并非统一应用干预,而是使用由干预相关子空间能量E_h确定的逐头系数g_h*来混合修复后的记忆和原始记忆,在推理时无需显式的危害分类器。在LLaVA-OneVision-7B、Chameleon-7B、Qwen2-VL-7B和InternVL2-4B上,SPARK在保持通用能力的同时降低了多模态攻击成功率。在LLaVA-OneVision-7B上,仅图像越狱攻击成功率降至4.7%,而MMMU保持在未防御模型0.6分以内(47.8对48.4),语言质量接近基线。在MM-SafetyBench上,攻击成功率从39.2%降至12.4%。即使在白盒自适应联合提示-图像攻击下,攻击成功率也被限制在20.3%,而未防御模型为54.6%。这些结果表明,通过在预填充阶段选择性地修复干预相关的KV子空间,可以大幅缓解多模态越狱行为,尤其是当有害证据由视觉模态携带时。
英文摘要
Vision-language models (VLMs) remain vulnerable to jailbreaks that distribute harmful intent across text and images, making unimodal safety mechanisms insufficient. We investigate whether this vulnerability can be mitigated directly in the multimodal key-value (KV) memory formed during prefill, without modifying model parameters at inference time. We introduce SPARK, a two-stage framework for targeted KV-memory repair. Stage 1 uses a disposable diagnostic adapter to identify harm-associated directions in multimodal key and value representations. Stage 2 projects out these directions, learns a lightweight residual repair, and anchors repaired keys with an image-structural prior to preserve visual grounding. Rather than applying the intervention uniformly, SPARK mixes repaired and original memory using a head-wise coefficient g_h* determined by intervention-relevant subspace energy E_h, requiring no explicit harm classifier at inference. Across LLaVA-OneVision-7B, Chameleon-7B, Qwen2-VL-7B, and InternVL2-4B, SPARK reduces multimodal attack success while preserving general capability. On LLaVA-OneVision-7B, image-only jailbreak attack success falls to 4.7%, while MMMU remains within 0.6 points of the undefended model (47.8 vs. 48.4) with near-baseline language quality. On MM-SafetyBench, attack success decreases from 39.2% to 12.4%. Even under white-box adaptive joint prompt-image attacks, attack success is limited to 20.3%, compared with 54.6% for the undefended model. These results suggest that multimodal jailbreak behavior can be substantially mitigated by selectively repairing intervention-relevant KV subspaces at prefill, particularly when harmful evidence is carried by the visual modality.
Comments23 pages (including appendix), 5 figures. Preprint