发表机构
The Chinese University of Hong Kong; IBM Research; University of Zagreb; Radboud University(香港中文大学; IBM研究院; 萨格勒布大学; 拉德堡德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SIR,一种针对操作系统层面计算使用智能体的黑盒间接提示注入攻击,通过迭代反馈循环提炼策略,使攻击成功率显著提升且可跨模型迁移。
AI 中文摘要
计算使用智能体(CUA)是一种视觉-语言模型,可感知屏幕并通过鼠标、键盘和终端在真实操作系统上执行操作,目前正越来越多地被部署用于自动化日常数字任务。由于这类智能体在运行时可能接触到不可信内容,它们易受到间接提示注入(IPI)攻击,攻击者会在智能体将读取的内容中植入指令,引导其执行违背用户意图的操作。现有的CUA安全基准测试采用人工编写的固定注入指令,可能低估自适应攻击者带来的风险。本文提出SIR,一种黑盒IPI攻击方法:(i)从一个小型可复用原则库中组合以自然语言表述的隐蔽注入指令;(ii)将组合过程包装为迭代反馈循环,该循环会诊断受害者的失败轨迹,并将绕过防御的方法提炼为新的命名策略,在不同任务间复用。与此前针对网页智能体的红队测试不同,本文针对操作系统层面的CUA,采用完全确定性的判定依据对攻击进行评分,通过检查文件系统、服务和权限状态而非大语言模型(LLM)裁判来完成评估。实验中,本文评估了三个前沿CUA。结合原则与反馈的组合方法使攻击成功率较人工编写的基准显著提升,例如Claude Opus 4.8上从4%升至24%,Gemini 3.5 Flash上从0%升至28%,同时仍能完成良性任务。针对一个模型发现的原则可进一步迁移至不同架构,无需额外反馈。
英文摘要
Computer-use agents (CUAs) are agents powered by vision-language models (VLMs) that perceive a screen and operate an operating system through mouse, keyboard, and terminal interactions to automate everyday digital tasks. Their exposure to untrusted content creates a risk of indirect prompt injection (IPI), where an adversary embeds instructions in content the agent reads to redirect it toward actions that violate the user's intent. Evaluations based on fixed, hand-written injections may underestimate the risk posed by adaptive adversaries. We present SIR, a black-box self-improving IPI framework that (i) composes task-specific injections from a small library of reusable red-teaming principles stated in plain language and (ii) uses an iterative feedback loop to analyze unsuccessful attack trajectories and distill new, named principles that are retained in a shared library and reused across tasks. We target operating-system-level compromise and evaluate outcomes through deterministic checks on filesystem, service, and permission state rather than an LLM judge. An attack counts as successful only when both the adversarial objective and the benign user task are completed in the same execution. We evaluate three frontier CUAs, allowing SIR up to 10 attack attempts per case. It achieves joint attack success rates of 24% on Claude Opus 4.8 and 28% on Gemini 3.5 Flash, compared with 4% and 0%, respectively, for the benchmark's fixed, hand-written injection. The red-teaming principles discovered against one victim also improve attacks against other victims, including a different model family, without additional feedback.