arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17050cs.CVcs.AI

EvoGUI:用于GUI状态转换理解的进化感知基准测试

EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yaohan Yang, Minglei Shi, Borui Zhang, Jie Zhou, Jiwen Lu

AI总结:

研究GUI状态转换理解,提出EvoGUI诊断框架,将GUI轨迹转化为视觉问答探针,无需额外注释。通过Mind2Web和WebLINX实例化EvoGUI-Bench并零样本评估28种模型配置,结果显示其可补充端到端评估,揭示理解提升空间。

AI中文摘要:

GUI智能体必须推断动作如何改变界面状态,但端到端成功率将这种能力与感知、基础、规划和恢复能力纠缠在一起。我们引入了EvoGUI,这是一个诊断框架,它将标准化的GUI轨迹转换为三个互补的视觉问答探针:时间排序、反向动作/值预测和对比性单步后继判别。它们的标签来自轨迹顺序和记录的动作,在轨迹标准化后无需额外的任务标签注释。我们从Mind2Web和WebLINX实例化了EvoGUI-Bench,在120个域中产生了3000个实例,并对28种视觉语言模型配置进行了零样本评估。最强的模型仅达到60.4 EvoGain,而模型规模和GUI专业化并不能可靠地预测性能。这些结果将EvoGUI-Bench确立为端到端GUI智能体评估的可扩展诊断补充,同时揭示了状态转换理解方面的巨大提升空间。源代码可在这个https URL上公开获取。

英文摘要:

GUI agents must reason about how actions transform interface states, but end-to-end success rates entangle this ability with perception, grounding, planning, and recovery. We introduce EvoGUI, a diagnostic framework that converts normalized GUI trajectories into three complementary visual question answering probes: temporal ordering, inverse action/value prediction, and contrastive one-step successor discrimination. Their labels are derived from trajectory order and logged actions, requiring no additional task-label annotation after trajectory normalization. We instantiate EvoGUI-Bench from Mind2Web and WebLINX, yielding 3,000 instances across 120 domains, and evaluate 28 vision-language model configurations zero-shot. The strongest model reaches only 60.4 EvoGain, while model scale and GUI specialization do not reliably predict performance. These results establish EvoGUI-Bench as a scalable diagnostic complement to end-to-end GUI-agent evaluation while exposing substantial headroom in state-transition understanding. The source code is publicly available at https://github.com/Yyhhh6/EvoGUI.

↑