AI 中文总结
研究针对可穿戴设备上视觉语言模型的物理提示注入攻击,利用真实环境照片分析出6种威胁向量及对12个模型的影响,攻击成功率高,还提出基于掩码的外部过滤器和基于语义向量的内部检测器两种防御策略。
AI 中文摘要
视觉语言模型(VLM)正迅速部署在如智能眼镜等面向人的可穿戴设备上,以实现多模态感知和人工智能辅助决策。先前研究已证明视觉提示注入数字图像输入的风险,但物理环境与可穿戴智能日益融合带来的独特安全挑战仍未得到充分探索。我们的工作刻画了物理环境中嵌入的恶意文本信息如何引入高优先级视觉通道进行间接提示注入,这种物理提示注入攻击不仅会扰乱基于VLM的可穿戴设备的正常任务,还会使模型产生不当输出。利用在200多个真实环境中从人工智能眼镜拍摄的照片,我们分析识别出6种代表性的物理注入提示威胁向量,并评估了它们对12个VLM模型的影响。结果表明这些攻击在关键任务中持续操纵模型输出,在模拟和真实环境中的攻击成功率分别高达96%和60%。我们还提出了两种针对性防御策略,有效降低了攻击成功率和安全影响。
英文摘要
Vision-Language Models (VLMs) are rapidly deployed on human-facing wearable devices such as smart glasses to enable multimodal perception and AI-assisted decision-making. While prior research has demonstrated the risks of visual prompt injection into digital image inputs of VLMs, the unique security challenges posed by the increasing integration between physical environments and wearable intelligence, such as those embodied in VLM-enabled AI glasses, remain underexplored. Toward understanding and modeling such threats, our work characterizes how malicious textual information embedded in physical environments introduces a high-priority visual channel for indirect prompt injection, where scene texts that hinder or evade human perception could hijack VLM models' behavior. Such \textit{Physical Prompt Injection Attacks} can not only disrupt normal tasks of VLM-enabled wearable devices, but also steer models to produce profane, biased, or even untruthful outputs. Using physically captured photos from AI glasses in over 200 real-world environments, our analysis identifies 6 representative threat vectors of physically injected prompts, and further evaluates their impacts on 12 VLM models. Results show that these attacks consistently manipulate model outputs across integrity- and safety-critical tasks, achieving attack success rates of up to 96\% and 60\% in simulated and real-world settings. Our analysis confirms that multiple models exhibit excessive blind trust in environmental text, ignoring the actual visual context and producing completely opposite summaries or directives. We further propose two targeted defense strategies, including a mask-based external filter and a semantic-vector-based internal detector, to effectively reduce the success rate and safety impact of these attacks.