多模态提示注入攻击对智能体AI框架的实验评估
An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks
浏览论文内容
中文总结 AI 辅助
本文提出MMPIBench基准,评估多模态提示注入对智能体框架的影响,发现攻击完成率低但模型差异显著,音频通道防御更弱。
中文摘要 AI 辅助
智能体AI框架使语言模型能够进行规划、保持记忆并调用可访问真实文件、邮件和服务的工具。这些智能体大多也能读取图像,这为攻击者提供了一种无需经过用户即可将文本注入智能体上下文的方式。我们提出了MMPIBench,一个可复现的基准测试,用于衡量注入后的后续影响。它通过六种视觉载体(OCR文本、叠加层、EXIF元数据、二维码、伪造界面及混合形式)传递一组固定的攻击,并记录每条注入指令在智能体中的传播距离,从感知、规划直至工具调用。在涵盖六个框架、五个基础模型、六种载体和四个攻击目标的720次运行中,攻击在约1%的运行中完成,但在12.8%的运行中被尝试,而这一差距几乎完全在规划步骤被弥合,即模型读取注入指令并拒绝执行。对于指令是否被执行,模型的作用远大于框架。一个模型从未尝试攻击,并在59.7%的运行中识别出注入,而另外两个模型在23.6%的运行中尝试攻击。随后,我们将基准扩展到音频,这是当前前沿模型接受的唯一其他原始感知通道。五个模型中只有两个能处理音频,六个框架中只有三个能传递音频,但在信号到达的地方,攻击在49%的单元中完成,对其中一个模型则为75%。因此,仅报告完成率会低估暴露风险,而视觉之外的感知通道虽然更窄,但防御力度也弱得多。
英文摘要
Agentic AI frameworks let a language model plan, keep memory, and call tools that reach real files, mail, and services. Most of these agents also read images, which gives an attacker a way to put text into the agent's context without going through the user. We present MMPIBench, a reproducible benchmark that measures what happens next. It delivers a fixed set of attacks through six visual carriers (OCR text, overlays, EXIF metadata, QR codes, fake interfaces, and hybrids) and records how far each injected instruction travels through the agent, from perception through planning to the tool call. Across 720 runs covering six frameworks, five foundation models, six carriers, and four attacker objectives, attacks complete in approximately 1% of runs but are attempted in 12.8%, and the gap is closed almost entirely at the planning step, where the model reads the injected instruction and declines to act on it. The model matters far more than the framework for whether an instruction is acted on. One model never attempts an attack and recognizes the injection in 59.7% of runs, while two others attempt in 23.6%. We then extend the benchmark to audio, the only other raw perceptual channel current frontier models accept. Only two of the five models ingest audio and only three of the six frameworks deliver it, but where the signal arrives the attack completes in 49% of cells, and in 75% for one model. Reporting completion alone therefore understates exposure, and perceptual channels beyond vision are narrower but much less defended.
发表机构
- California State Polytechnic University, Pomona(加州州立理工大学波莫纳分校)
机构由 AI 辅助整理,请以论文原文为准。