ComponentBench:诊断计算机使用智能体的组件级故障
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
浏览论文内容
中文总结 AI 辅助
本文提出ComponentBench基准及诊断流水线,评估7种计算机使用智能体在4种观测动作空间下的表现,发现观测动作空间会显著影响性能,且空间操作仍是智能体的挑战。
中文摘要 AI 辅助
当前对计算机使用智能体的评估分为长周期工作流基准测试和原子级GUI接地测试,这导致中间层存在评估不足的问题:即现实中以组件为中心的交互(例如切换一组按钮),这类交互足够短以便诊断,又足够丰富以捕捉现代界面的复杂程度。本文提出ComponentBench,这是一个用于在现代Web UI上对计算机使用智能体进行组件级评估的基准和诊断流水线。ComponentBench围绕一个与库无关的本体构建,包含97种规范UI组件,这些组件实例化为2910个经程序验证的任务,覆盖广泛使用的组件库,同时搭配经过清理的人类参考轨迹,可用于评估任务成功率和交互效率。除任务收集外,本文还引入了可扩展的流水线,用于在实施后审计实际的结构难度,并跨任务和组件族综合结构化故障分析。通过在四个观测和动作空间上评估七个模型——GPT-5.4、Gemini 3 Flash、GPT-5.4 mini、GPT-5 mini、Gemini 3.1 Flash-Lite、Qwen3-VL-235B和UI-TARS-1.5-7B,本文表明这些设计选择对性能有重大影响:在单个共享测试框架内,仅改变观测和动作空间,同一模型的任务成功率变化超过30%,例如GPT-5 mini在使用可访问性树观测时成功率为83.1%,而仅使用坐标的Pixel控制时降至48.9%;此外,即使是最快的配置,耗时也达到匹配人类参考的3.7倍,对人类来说微不足道的空间操作仍然是当前智能体面临的挑战。
英文摘要
Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an under-instrumented middle layer: realistic component-centered interactions (e.g., toggle a button set) that are short enough to diagnose and rich enough to capture the burdens of modern interfaces. We present ComponentBench, a benchmark and diagnostic pipeline for component-level evaluation of computer-use agents on modern web UIs. ComponentBench is organized around a library-agnostic ontology of 97 canonical UI components instantiated as 2,910 programmatically verified tasks across widely used component libraries, paired with cleaned human reference trajectories that enable evaluation of both task success and interaction efficiency. Beyond task collection, we introduce a scalable pipeline for auditing realized structural difficulty after implementation and synthesizing structured failure analyses across tasks and component families. Evaluating seven models -- GPT-5.4, Gemini 3 Flash, GPT-5.4 mini, GPT-5 mini, Gemini 3.1 Flash-Lite, Qwen3-VL-235B, and UI-TARS-1.5-7B -- across four observation and action spaces, we show that these design choices critically impact performance. Within a single shared harness, changing only the observation and action space shifts task success by more than 30% for the same model: GPT-5 mini falls from 83.1% with accessibility-tree observations to 48.9% with coordinate-only Pixel control. Moreover, even the fastest configuration takes 3.7x as long as the matched human reference, and spatial manipulations that are trivial for humans continue to challenge current agents.
发表机构
- Duke University(杜克大学)
- Amazon AGI SF Lab(亚马逊AGI旧金山实验室)
机构由 AI 辅助整理,请以论文原文为准。