发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SCOUT是一个两阶段智能体安全验证器,通过推理密集的规则生成与工具密集的证据收集协同,提升计算机操作智能体的安全检测性能,并在多个基准上取得领先结果。
AI 中文摘要
计算机操作智能体(CUA)虽然能够在日常和专业工作流程中完成计算机任务,但即使在良性的指令和环境条件下,也可能造成意外伤害。然而,检测此类伤害仍然具有挑战性。首先,这需要细致且任务特定的推理:仅由通用安全标准引导的验证器往往会忽略许多重要但细微的有害行为。其次,这需要主动调查:过去的轨迹截图显示了智能体做了什么,但并不总能显示环境中实际发生了什么变化,因此仅依赖截图的LLM-as-a-judge验证器可能无法确定行动的实际后果。为应对这些挑战,我们引入了SCOUT,一个两阶段的智能体安全验证器,它将推理密集的规则生成与工具密集的证据收集协同起来。首先,我们的SCOUT规则生成器对任务和智能体的轨迹进行广泛推理,以确定成功且安全执行应包含的内容,生成任务特定的完成规则和安全规则。然后,我们的SCOUT探测智能体遵循这些规则,与执行后的环境进行交互,并为最终的安全和完成判断收集基于实地的证据。我们在两个计算机使用安全基准上评估了我们的框架。在AutoElicit-Bench上,SCOUT实现了75.4的不安全F1分数和74.5的完成F1分数,优于LLM-as-a-judge验证器和朴素的工具使用验证器。SCOUT在OS-Blind上以76.4%的不安全检测准确率领先。测试时反思将AutoElicit-Bench上的最终不安全执行率从30.2%降低到17.2%。消融和分析表明,SCOUT中无工具的规则生成引发了更多的推理,并且对于跨验证器主干(尤其是非前沿主干)的安全检测至关重要。对编码任务的初步扩展表明,SCOUT可以支持计算机使用之外的安全验证。
英文摘要
Computer-use agents (CUAs), while capable of completing computer tasks in everyday and professional workflows, can cause unintended harm even under benign instructions and environments. However, detecting such harm remains challenging. First, it requires careful, task-specific reasoning: verifiers guided only by general safety criteria often overlook many important but subtle harmful behaviors. Second, it requires active investigation: past trajectory screenshots show what the agent did but not always what actually changed in the environment, so LLM-as-a-judge verifiers that rely on screenshots alone may be unable to determine the actual consequences of actions. To address these challenges, we introduce SCOUT, a two-stage agentic safety verifier that synergizes reasoning-intensive rubric generation with tool-intensive evidence gathering. First, our SCOUT rubric generator extensively reasons over the task and the agent's trajectory to determine what successful and safe execution should entail, generating task-specific completion and safety rubrics. Then, our SCOUT probing agent follows these rubrics to interact with the post-execution environment and collect grounded evidence for final safety and completion judgments. We evaluate our framework on two computer-use safety benchmarks. On AutoElicit-Bench, SCOUT achieves 75.4 unsafe F1 and 74.5 completion F1, outperforming LLM-as-a-judge verifiers and naive tool-use verifiers. SCOUT leads on OS-Blind with 76.4% unsafe detection accuracy. Test-time reflection reduces final unsafe execution rates from 30.2% to 17.2% on AutoElicit-Bench. Ablations and analysis show that tool-free rubric generation in SCOUT elicits substantially more reasoning and is crucial for safety detection across verifier backbones, especially non-frontier ones. A preliminary extension to coding tasks shows that SCOUT can support safety verification beyond computer-use.