Secure-CUA:控制计算机使用代理中的不可信影响
Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents
- University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
- Google(谷歌)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Secure-CUA通过操作事务和隔离查询模型,在计算机使用代理中控制不可信内容影响,实现安全执行,并在WebArena任务中达到53.55%的成功率。
AI中文摘要:
计算机使用代理(CUA)通过观察图形界面并发出点击和按键等命令,在跨应用程序(如桌面、移动应用和网页浏览器)中执行任务。这些界面将可信控件和内容与合法任务所需的不可信内容相结合。控制这些不可信内容的对手可以嵌入指令或误导性视觉线索,以改变代理的预期操作或将其命令重定向到错误的界面目标。我们为代理的决策及其通过GUI命令的执行形式化了安全要求。在理想的执行模型中,我们表明在每一步强制执行这两个要求可以保护执行轨迹。我们在Secure-CUA中实例化了该模型,这是我们的安全CUA执行系统。其关键思想是在访问不可信内容之前,承诺一个显式的逐操作程序,称为“操作事务”。每个事务固定其对不可信内容的查询以及对其响应的允许用途。系统屏蔽不可信区域,并使用隔离查询模型评估每个事务以产生下一个操作,从而回答其查询。然后,它使用屏蔽后的界面定位预期的界面目标。在该模型的假设下,Secure-CUA在设计上是安全的,而每一步生成新事务有助于通过适应不断变化的界面来保持高任务效用。我们在良性条件下使用三个前沿模型在400个WebArena任务上评估了Secure-CUA,跨5个种子,产生了6,000条执行轨迹。Secure-CUA实现了53.55%的平均任务成功率,而Vanilla-CUA为55.12%,CaMeL-CUA为13.17%。
英文摘要:
Computer-use agents (CUAs) perform tasks across applications (such as desktops, mobile apps, and web browsers) by observing graphical interfaces and issuing commands such as clicks and keystrokes. These interfaces combine trusted controls and content with untrusted content needed for legitimate tasks. An adversary controlling this untrusted content can embed instructions or misleading visual cues to change the agent's intended action or redirect its commands to the wrong interface target. We formalize security requirements for both the agent's decisions and their execution through GUI commands. In an ideal execution model, we show that enforcing both requirements at each step protects execution traces. We instantiate this model in Secure-CUA, our system for secure CUA execution. Its key idea is to commit to an explicit per-action program, called an $\textit{action transaction}$, before accessing untrusted content. Each transaction fixes its queries to untrusted content and the permitted uses of their responses. The system masks untrusted regions and evaluates each transaction to produce the next action, using an isolated query model to answer its queries. It then locates the intended interface target using the masked interface. Under the model's assumptions, Secure-CUA is secure by design, while generating a new transaction at each step helps maintain high task utility by adapting to changing interfaces. We evaluate Secure-CUA under benign conditions on 400 WebArena tasks using three frontier models across $5$ seeds, yielding $6,000$ execution traces. Secure-CUA achieves an average task success rate of $53.55\%$, compared with $55.12\%$ for Vanilla-CUA and $13.17\%$ for CaMeL-CUA.