发表机构
MATS, EPFL(MATS,洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Control OSWorld,一个结合318个正常任务与81个有害任务的GUI智能体控制评估环境,并研究监控器检测恶意行为,发现完整轨迹下监控效果佳,但逐步评分时性能下降,文本信息比截图更重要。
AI 中文摘要
通过图形用户界面(GUI)操作计算机的AI智能体正被广泛部署。AI控制研究旨在防止AI系统在即使其与人类意图不一致并主动试图造成伤害的情况下造成危害。迄今为止,大多数控制研究集中于编码智能体,对计算机使用场景的探索相对较少。我们引入了Control OSWorld,这是一个控制评估环境,将OSWorld中的318个任务与81个有害侧任务(例如,窃取私人文件)配对,要求智能体在未被发现的情况下完成这些有害任务。我们利用Control OSWorld创建并研究了用于检测与GUI交互的恶意智能体的控制监控器。我们发现,当监控器能够看到完整轨迹时,一个较弱的监控器可以可靠地区分由更强智能体产生的诚实轨迹与恶意轨迹(在3%的假阳性率下,召回率达到97%)。然而,当监控器必须在每一步执行前进行评分时,由于每个动作的评估上下文较少,低假阳性率下的召回率会下降。监控器的性能还取决于其可见内容:访问智能体的可见文本至关重要,而屏幕截图提供的额外价值很小。
英文摘要
AI agents that operate a computer through its graphical user interface (GUI) are being widely deployed. AI control studies how to prevent an AI system from causing harm even if it is misaligned and actively trying to do so. Most control research to date has focused on coding agents, leaving computer use largely unexplored. We introduce Control OSWorld, a control evaluation that pairs 318 tasks from OSWorld with 81 harmful side tasks (e.g., exfiltrating a private file) that an agent must complete without being caught. We use Control OSWorld to create and study control monitors that detect malicious agents interacting with a GUI. We find that a weaker monitor can reliably distinguish honest from malicious trajectories produced by a more capable agent when it sees the full trajectory (97% recall at 3% false positive rate). However, when the monitor has to score each step before it is executed, recall drops at low false positive rates because each action is judged with less context. Monitor performance also depends on what the monitor can see: access to the agent's visible text is critical, while screenshots provide little additional value.