面向多模态大语言模型的可迁移越狱攻击的文本锚定语义扰动
Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models
查看机构详情
- Harbin Institute of Technology(哈尔滨工业大学)
- Pengcheng Laboratory(鹏城实验室)
- Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究针对多模态大语言模型易受越狱攻击的问题,提出TA-SPA黑盒越狱框架,通过文本锚定语义分解与语义保留增强实现可迁移攻击,在商业模型及防御下表现出竞争力,推动表征级安全对齐研究。
中文摘要 AI 辅助
多模态大语言模型(Multimodal Large Language Models, MLLMs)在视觉-语言交互领域取得了显著进展,但其安全对齐仍易受到越狱攻击。一个关键挑战在于,在文本空间中学习到的安全行为无法可靠地迁移到融合的跨模态表征,使得多模态输入可通过潜在语义线索被利用。我们提出文本锚定语义扰动攻击(Text-Anchored Semantic Perturbation Attack, TA-SPA),这是一种黑盒越狱框架,可在文本锚定语义空间中优化可迁移扰动。TA-SPA整合了文本锚定语义分解(Text-Anchored Semantic Factorization, TASF)与语义保留增强(Semantic-Preserving Augmentation, SPA):TASF鼓励将跨模态语义因素与模态特定残差分离,SPA则在保留语义一致性的同时多样化有害目标锚点。实验表明,该攻击具有较强的有效性,且可迁移至商业多模态大语言模型,在代表性防御下表现出竞争力。额外的控制与探测实验支持了预期的分解效果(不暗示完全解耦),为超越输入级过滤的表征级安全对齐提供了思路。
英文摘要
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual space does not reliably transfer to fused cross-modal representations, leaving multimodal inputs exploitable through latent semantic cues. We propose Text-Anchored Semantic Perturbation Attack (TA-SPA), a black-box jailbreak framework that optimizes transferable perturbations in a text-anchored semantic space. TA-SPA integrates Text-Anchored Semantic Factorization (TASF), which encourages the separation of cross-modal semantic factors from modality-specific residuals, with Semantic-Preserving Augmentation (SPA), which diversifies harmful target anchors while preserving semantic consistency. Experiments show strong attack effectiveness and transfer to commercial MLLMs, with competitive performance under representative defenses. Additional controls and probing support the intended factorization without implying perfect disentanglement, motivating representation-level safety alignment beyond input-level filtering.