发表机构
ETH Zurich; Leiden University(苏黎世联邦理工学院; 莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现多模态模型对文本指令更敏感,利用此模态差距提出无需训练的防御方法Pictionary,将不可信内容渲染为图像以降低提示注入攻击成功率,并验证其有效性。
AI 中文摘要
大型语言模型容易受到提示注入攻击,其中第三方对抗性内容可以劫持模型的行为。在本文中,我们研究了对抗性数据的输入模态所扮演的角色,并识别出一个系统性不对称:多模态LLM在指令以文本形式出现时,比通过非文本渠道(如图像)传递相同指令时,更有可能遵循对抗性指令。我们假设这种模态差距源于以文本为中心的指令微调,这种微调教会模型服从文本指令,而将其他模态主要视为要解析或描述的内容。然后,我们展示了如何将这种差距转化为一种无需训练的防御,方法是在所有不可信载荷到达模型之前,将其渲染为排版图像(或音频)。在十个模型和两个提示注入基准(DirectInject和AgentDojo)上,我们表明我们的防御方法Pictionary持续降低攻击成功率,即使面对最强的自适应攻击和人类红队,同时基本保持良性实用性。我们进一步表明,对图像渲染指令进行良性微调会侵蚀模态差距,将其追溯到以文本为中心的指令微调分布。
英文摘要
Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to obey textual instructions while treating other modalities mainly as content to parse or describe. We then demonstrate how this gap can be turned into a training-free defense, by rendering all untrusted payloads as typographic images (or audio) before they reach the model. Across ten models and two prompt injection benchmarks (DirectInject and AgentDojo) we show that our defense Pictionary consistently reduces attack success rates even against the strongest adaptive attacks and human red teamers, while largely preserving benign utility. We further show that benign fine-tuning on image-rendered instructions erodes the modality gap, tracing it to the text-centric instruction-tuning distribution.