arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

保护计算机使用代理免受分支转向攻击

Securing Computer-Use Agents Against Branch Steering Attacks

Giulio Zingrillo, Hanna Foerster, Ilia Shumailov, Yiren Zhao, Robert Mullins

arXiv 2610.03089首次发表:更新:

发表机构

ETH Zurich; University of Cambridge; AI Sequrity Company; Imperial College London(苏黎世联邦理工学院; 剑桥大学; AI安全公司; 帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对计算机使用代理的分支转向攻击,提出COBRA架构,通过可信分支计划与能力约束,将攻击成功率降至0%,同时保留97%的良性效用。

AI 中文摘要

现代计算机使用代理(CUA)直接与图形用户界面交互并执行第三方网络工具,这使得它们暴露于每个渲染页面和工具响应中的间接提示注入。虽然双LLM模式是提供正式安全保证的主要系统级架构——使用隔离的计划LLM(P-LLM)在处理不可信输入之前固定执行路径,并通过隔离LLM(Q-LLM)处理——但这些保证在图形环境中失效。由于CUA交互本质上是动态的,计划不能保持数据无关;它们必须根据预期的运行时网络内容进行分支,覆盖代理可能遇到的所有可能情况。这使得代理面临分支转向攻击,攻击者构造不可信数据以胁迫CUA沿着危险的、预先批准的分支执行,而无需注入显式指令。我们系统地研究了分支转向攻击,并引入了STEER-Bench(涵盖9个领域的101个任务),表明对标准CUA(94.4%)和普通双LLM CUA(89.5%)的攻击成功率都很高。然后,我们提出了COBRA,一种将可信分支计划与提前能力约束相结合的架构,严格限制每个分支可能执行的参数和目的地。在STEER-Bench上,COBRA将攻击成功率降低到0%,同时保留了97%的良性效用。

英文摘要

Modern Computer Use Agents (CUAs) directly interact with graphical user interfaces and execute third-party web tools, exposing them to indirect prompt injection across every rendered page and tool response. While the Dual-LLM pattern is the primary system-level architecture offering formal security guarantees - using an isolated Planner LLM (P-LLM) to fix execution paths before processing untrusted inputs via a Quarantined LLM (Q-LLM) - these guarantees break down in graphical environments. Because CUA interaction is inherently dynamic, plans cannot remain data-independent; they must branch based on anticipated runtime web content - covering all possible cases the agent may encounter. This exposes agents to branch steering attacks, where an adversary crafts untrusted data to coerce a CUA down a hazardous, pre-approved branch without injecting explicit instructions. We systematically study branch steering attacks and introduce STEER-Bench (101 tasks across 9 domains), showing high attack success against both standard (94.4%) and vanilla Dual-LLM (89.5%) CUAs. We then propose COBRA, an architecture that pairs trusted branching plans with ahead-of-time capability constraints, strictly bounding the parameters and destinations each branch may execute. On STEER-Bench, COBRA reduces attack success to 0% while retaining 97% benign utility.

Comments16 pages, including 2 figures. To be presented at the "Agents in the Wild" Workshop at the NeurIPS 2026 Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑