ADeptS-Bench:衡量跨设备计算机使用智能体的可信赖性
ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
浏览论文内容
中文总结 AI 辅助
ADeptS-Bench是基于ADEPTS框架的双流基准,评估7个计算机使用智能体的跨设备可信赖性,发现其存在安全与歧义处理缺陷,将发布相关数据与工具。
中文摘要 AI 辅助
计算机使用智能体(CUAs)越来越多地被部署以代表用户导航移动和桌面应用程序,但目前尚无基准能够全面评估它们在处理模糊指令时能否与视觉界面安全交互。我们引入了基于ADEPTS能力框架和普通用户研究的双流可信赖性基准ADeptS-Bench。安全流提供嵌入视觉界面威胁的成对良性/恶意任务;歧义流评估智能体在意图模糊时是否会寻求澄清。对7个模型的评估显示,没有模型能在保持攻击成功率低于30%的同时,任务成功率始终超过80%;所有模型都会毫不犹豫地点击25000美元订单的“结账”按钮,且没有一个能检测到“出厂重置”按钮被错误标记为“优化”。消融实验揭示了三种不同的安全架构:依赖工具型(无拒绝工具时ASR提升21-23个百分点)、部分依赖工具型(提升10-11个百分点)、无机制型(无变化)。在歧义处理中,所有模型都高估了后果严重性,与安全领域观察到的过度拒绝偏差相似。我们将在发表后发布所有数据、评估代码及分析工具。
英文摘要
Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on behalf of users, yet no benchmark comprehensively evaluates whether they can safely interact with visual interfaces while handling ambiguous instructions. We introduce ADeptS-Bench, a dual-stream trustworthiness benchmark, grounded in the ADEPTS capability framework and general population user studies. The Safety stream provides paired benign/malicious tasks with threats embedded in the visual interface. The Disambiguation stream evaluates whether agents seek clarification when intent is ambiguous. Evaluating seven models reveals that no model consistently exceeds 80% task success while staying below 30% attack success; every model clicks "Checkout" on a $25K order without hesitation, and none detects that a "factory reset" button is mislabeled as "Optimize." An ablation reveals three distinct safety architectures: tool-dependent (ASR +21-23pp without refusal tool), partially tool-dependent (+10-11pp), and no mechanism (unchanged). In disambiguation, all models overestimate consequence severity, mirroring the over-refusal bias observed in safety. We release all data, evaluation code, and analysis tools upon publication.
发表机构
- FAIR at Meta(Meta FAIR实验室)
机构由 AI 辅助整理,请以论文原文为准。