arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跟我重复:黑盒自适应视觉提示注入

Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection

Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri, Raluca Ada Popa, David Wagner, Arman Zharmagambetov

arXiv 2609.04533首次发表:更新:

发表机构

UC Berkeley; FAIR at Meta(加州大学伯克利分校; Meta 基础人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Repeat-After-Me黑盒自适应视觉提示注入攻击,在Qwen3.6-27B、GPT-5.5等VLM上实现较高攻击成功率,可用于泄露隐私或发起恶意工具调用,在真实OpenClaw智能体中也能生效。

AI 中文摘要

提示注入被广泛认为是与不可信外部数据(如网站、文档和电子邮件)交互的AI智能体的主要安全威胁。先前研究表明,在文本领域,黑盒提示注入可达到接近完美的攻击成功率(ASR)。然而在图像领域,现有视觉提示注入方法在攻击前沿商用视觉语言模型(VLM)以实现实质性有害行为时效果要差得多。实现此类输出难度较大,因为它需要一个长且符合格式的目标字符串,例如带有精确函数名和参数的可解析原生工具调用。本文提出Repeat-After-Me,一种黑盒自适应视觉提示注入攻击,可泄露个人身份信息或发起恶意工具调用。在包括Qwen3.6-27B和GPT-5.5在内的开源权重及商用前沿VLM上,在良性用户提示与注入任务语义无关且未口头授权的现实场景中,我们的方法分别达到超过80%和47%的ASR。评估显示,在一个替代模型上优化的注入,在两个商用受害模型上保留了43%-46%的原始ASR;跨样本可迁移性在这两个模型上保留了64%-66%的原始ASR。我们在真实世界的OpenClaw智能体中测试了该攻击:在默认的OpenClaw Discord部署中,不可信用户可使用经最小程度注入的图像覆盖该链接,从而实现远程代码执行和秘密数据 exfiltration 等敏感行为。我们证明了该新攻击向量在自适应文本提示注入失效的场景下依然有效,并讨论了潜在防御措施。

英文摘要

Prompt injection is widely recognized as a major security threat to AI agents that interact with untrusted external data, such as websites, documents, and emails. Prior work has shown that, in the text domain, black-box prompt injection can achieve near-perfect attack success rates (ASRs). In the image domain, however, existing visual prompt injection methods are substantially less effective in attacking frontier commercial VLMs for materially harmful behavior. Achieving such outputs is hard because it requires a long and/or format-compliant target string, such as a precise, parseable native tool call with exact function names and arguments. We present Repeat-After-Me, a black-box adaptive visual prompt injection attack that can reveal personally identifiable information or make malicious tool calls. Across both open-weight and commercial frontier VLMs, including Qwen3.6-27B and GPT-5.5, our method achieves ASRs exceeding 82% and 47%, respectively, under a realistic setting in which the benign user prompt is semantically unrelated to the injected task and does not verbally authorize it. Our optimized injection has non-trivial attack transferability across commercial VLMs and benign samples. We show our attack works in cases where adaptive textual prompt injection fails. In a real-world OpenClaw agent connected to Discord, an untrusted user can use a minimally injected image from our attack to overwrite TOOLS.md, enabling future sensitive behaviors like remote code execution and secret exfiltration. We discuss potential defenses.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑