发表机构
University of Illinois Urbana-Champaign; Yale University(伊利诺伊大学厄巴纳-香槟分校; 耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出CUADebug框架,含错误分类、CUAErrorBench基准与CUADebugger工具,可诊断修复计算机使用智能体故障,提升任务完成与重新执行成功率,证明其能提供可操作修复信号。
AI 中文摘要
计算机使用智能体(CUA)通过截图、鼠标和键盘操作以及有状态的UI反馈来操作真实的桌面和网页界面,但其故障仍难以诊断和修复。与仅处理文本的智能体不同,CUA故障源于视觉感知、空间定位、低级交互、任务推理和环境动态的耦合,使得调试成为一个独特的多模态因果定位问题。我们提出了CUADebug,一个用于诊断和修复CUA故障的框架。CUADebug包含CUA特定的错误分类、人类标注的OSWorld故障基准CUAErrorBench,以及工具增强调试器CUADebugger。CUADebugger并非仅对完整轨迹进行一次提示,而是通过成对的前后截图和操作轨迹主动检查可疑步骤,随后提交结构化诊断,包含根本原因步骤、错误类型、定位证据和重新执行的纠正策略。对204条失败轨迹的人类标注显示,任务推理与控制是最大的故障类别(110/204),其次是感知(36)、定位/交互(25)、外部/系统(13),以及20个OSWorld不可行任务案例的其他类别。在主要的Claude-agent划分上,CUADebugger结合Gemini 2.5 Pro将联合子类型与步骤诊断准确率从11.2%提升至19.6%,且在不同调试器主干上均有一致提升。在单次重新执行评估中,基于根本原因分析(RCA)的条件相比仅基于历史的延续,实现了更高的任务完成率(机器RCA为28.47%,我们的方法为29.90%,而后者为13.89%);在持续重新执行中,我们的方法将成功率从12.2%提升至25.86%,而人类专家指导达到29.21%。这些结果表明,CUA根本原因诊断可提供可操作的修复信号,而非仅事后解释。
英文摘要
Computer-use agents (CUAs) interact with graphical interfaces through screenshots and low-level mouse and keyboard actions, yet the causal error may precede the terminal failure. We present CUADebug, a framework for localizing root causes in CUA trajectories and guiding re-execution. CUADebug includes a five-category, 30-subtype taxonomy; CUAErrorBench, a benchmark of 204 failed OSWorld trajectories with human root-cause annotations; and CUADebugger, a ReAct-style agent for root-cause analysis (RCA). CUADebugger iteratively selects trajectory steps, inspects paired before/after screenshots and action traces, and submits a structured diagnosis containing the causal step, taxonomy label, grounded evidence, and correction. CUADebugger performs RCA without per-trajectory human intervention; human annotations are used to evaluate RCA predictions and, in controlled re-rollout comparisons, to fix restart points. Task reasoning and control is the largest annotated failure category (110/204). CUADebugger improves L2 and Tag+Step Exact across three debugger backbones on the Claude-agent split; with Gemini 2.5 Pro, Tag+Step Exact rises from 11.1% to 19.4%. Single re-execution improves failure recovery from 13.89% to 29.86% (overall 61.77% to 68.14%); controlled continual re-execution improves it from 12.50% to 25.69% (overall 61.22% to 66.48%). Project page: cuadebug.github.io.
Comments23 pages, 10 figures, 6 tables