arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34113cs.AI

GUITAR:通过状态转换对GUI代理进行结构化失败诊断

GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions

Shaoqing Zhang, Kehai Chen, Xuefeng Bai, Zhuosheng Zhang, Pengfei Zhang, Yang Xiang, Min Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

GUITAR提出状态中心诊断框架,通过状态转换图分析GUI代理失败,定位瓶颈,提升成功率2.8%。

中文摘要 AI 辅助

理解图形用户界面(GUI)代理在何处以及为何失败,对于构建更可靠的系统至关重要,然而当前的评估依赖于步骤准确率这一指标,该指标将每个屏幕独立对待,忽视了GUI环境的底层结构。这导致两个关键的盲点:(1)功能等效的屏幕被孤立评估,掩盖了共享屏幕上的系统性失败模式;(2)长尾GUI分布使得标准指标下罕见但关键屏幕上的失败变得不可见。为解决这些问题,我们提出了GUITAR,一个以状态为中心的诊断框架,通过将视觉多样的屏幕映射到共享功能状态,使用状态转换图(STG)对状态和转换进行结构化失败分析。在来自AndroidControl和Mind2Web的8个代理和6个任务中,GUITAR揭示了60.4%的失败发生在20%的状态中,将错误定位到一小部分瓶颈。瓶颈定向指导将成功率(SR)提高了2.8%,并在三倍轨迹留出评估下,在7个代理中保持了1.88%的平均增益,且完全自动生成STG。这些发现证明了在评估的移动和网页任务中,结构感知评估的诊断和可操作价值。代码可在该https URL获取。

英文摘要

Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Across 8 agents and 6 tasks from AndroidControl and Mind2Web, GUITAR reveals that 60.4\% of failures occur in 20\% of states, localizing errors to a small set of bottlenecks. Bottleneck-targeted guidance improves SR by 2.8\% and retains a 1.88\% average gain across 7 agents under three-fold trajectory-held-out evaluation with fully automatic STGs. These findings demonstrate the diagnostic and actionable value of structure-aware evaluation within the evaluated mobile and web tasks. Code is available at https://github.com/sqzhang-lazy/GUITAR

发表机构

  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
  • Pengcheng Laboratory(鹏城实验室)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑