发表机构
Tsinghua University; Harbin Engineering University(清华大学; 哈尔滨工程大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
UnifiedAttack通过协同劫持框架(ICR和CPI)评估大型多模态模型在协同有害图文生成中的安全性,发现统一模型的结构性帮助性和逻辑一致性可被系统性武器化,亟需逻辑感知防御。
AI 中文摘要
随着大型多模态模型(LMMs)向原生统一架构过渡,评估其在协同有害图文生成任务中的安全性成为一个关键挑战。与单模态威胁不同,协同风险出现在文本和图像模态协调产生危害时,其危害程度显著超过各模态单独作用之和。我们引入了UnifiedAttack,这是一个新颖的基准,旨在通过关注跨模态协同所实现的危害增益来评估LMM在协作场景中的安全性。该基准包含了根据其多模态潜力筛选的样本,以及一组新颖的合成虚假信息查询。为了验证已识别的漏洞,我们提出了一种协同劫持框架,包括上下文重新包装(ICR)和认知规划注入(CPI)。ICR利用少样本学习将对抗意图包裹在良性虚拟外壳中,以使安全过滤器脱敏,而CPI通过强制执行先规划后执行的范式来劫持推理路径。通过迫使系统承诺一个中立的逻辑计划,我们利用其内部一致性驱动来诱导同步生成有害的多模态内容。对最先进架构的广泛评估表明,UnifiedAttack始终能够绕过现代对齐机制。我们的发现揭示了统一模型的结构性帮助性和逻辑一致性可以被系统性地武器化,突显了在协同生成任务中迫切需要逻辑感知防御。代码可在该https URL获取。
英文摘要
As Large Multimodal Models (LMMs) transition toward natively unified architectures, evaluating their safety in synergistic harmful image-text generation tasks becomes a critical challenge. Unlike unimodal threats, synergistic risks emerge when text and image modalities are coordinated to produce harm that significantly exceeds their individual components. We introduce UnifiedAttack, a novel benchmark designed to evaluate LMM safety in collaborative scenarios by focusing on the harmfulness gain achieved through cross-modal synergy. The benchmark incorporates samples filtered for their multimodal potential alongside a novel subset of synthesized disinformation queries. To verify identified vulnerabilities, we propose a synergistic hijacking framework featuring In-Context Reskinning (ICR) and Cognitive Planning Injection (CPI). ICR utilizes few-shot learning to wrap adversarial intent in benign virtual shells to desensitize safety filters, while CPI hijacks the reasoning path by enforcing a plan-then-execute paradigm. By compelling the system to commit to a neutral logical plan, we exploit its internal drive for consistency to induce the synchronized generation of harmful multimodal content. Extensive evaluations on state-of-the-art architectures demonstrate that UnifiedAttack consistently bypasses modern alignment. Our findings reveal that the structural helpfulness and logical coherence of unified models can be systematically weaponized, highlighting the urgent need for logic-aware defenses in synergistic generation tasks. Code is available at https://github.com/bingjunluo/UnifiedAttack .