用一张纸劫持机器人:对VLM控制机器人的物理提示注入的系统研究
Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots
浏览论文内容
中文总结 AI 辅助
该研究针对VLM控制的分拣机器人,系统分析了四类物理提示注入攻击的成功率,发现其存在显著脆弱性,且三种简单防御措施可大幅降低风险。
中文摘要 AI 辅助
视觉语言模型(VLM)越来越多地被部署为机器人系统中的规划器,它们将自然语言命令转换为基于视觉场景理解的可执行动作。感知与指令遵循之间的这种紧密耦合引入了新的攻击面:放置在机器人视野内的对抗性文本可作为对VLM推理栈的间接提示注入。我们针对VLM控制的分拣任务开展了物理提示注入攻击的系统研究,提出了四类分类法:间接标识、任务重定义、权威冒充、冲突注入,实例化为包含20个攻击提示的基准,在三种物理场景布局和三种命令形式(目的地特异性和规则明确性存在差异)下进行评估。针对三个前沿VLM(GPT-4o、Gemini 2.5 Flash、Qwen3-VL-32B)开展的5670次试验显示,攻击成功率分别为27.0%、29.4%和5.0%,其中权威冒充和否定攻击可在所有三个模型间迁移。对推理轨迹的分析表明,成功的妥协几乎总是有意识的(99.9%的确认率),且模型通过结构不同的机制进行防御:Gemini采用明确拒绝,GPT-4o采用感知注意力缺失。我们评估了三种简单的防御措施:基于提示的防御(75%-100%有效,因模型而异)、两阶段验证(85%-100%)、预处理文本掩码(100%)。研究结果表明,VLM控制的操纵对人类可读的物理标识存在显著脆弱性,简单防御可大幅降低风险,但防御选择存在权衡。这些防御措施在我们的基准中保留了一般任务能力,但可能会损害需要读取场景内标签的任务。
英文摘要
Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding. This tight coupling between perception and instruction-following introduces a new attack surface: adversarial text placed within the robot's visual field can act as an indirect prompt injection into the VLM's reasoning stack. We present a systematic study of physical prompt injection attacks against VLM-controlled sorting, introducing a four-category taxonomy, indirect signage, task redefinition, authority impersonation, and conflict injection, instantiated as a benchmark of 20 attack prompts evaluated across three physical scene layouts and three command formulations that vary in destination specificity and rule explicitness. Across 5,670 trials on three frontier VLMs (GPT-4o, Gemini 2.5 Flash, Qwen3-VL-32B), attacks succeed at 27.0%, 29.4%, and 5.0% respectively, with authority-impersonating and negation attacks transferring across all three models. Analysis of reasoning traces reveals that successful compromise is nearly always conscious (99.9% acknowledgment rate), and that models defend through structurally different mechanisms, explicit rejection for Gemini, perceptual inattention for GPT-4o. We evaluate three simple mitigations: prompt-based defense (75-100% effective, model-dependent), two-stage verification (85-100%), and pre-processing text masking (100%). Our findings show that VLM-controlled manipulation is meaningfully vulnerable to human-readable physical signage, and that simple defenses substantially reduce risk, though defense choice involves trade-offs. The defenses preserve general task capabilities in our benchmark, but they may impair tasks that require reading in-scene labels.