arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00948cs.LGcs.AI

GUI-HARVEST:通过证据驱动的工具链进化实现自改进的GUI智能体

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Tianjin University(天津大学)
  • Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
  • East China Normal University(华东师范大学)
  • Shenzhen Research Institute of Big Data(深圳市大数据研究院)

机构由 AI 辅助整理,请以论文原文为准。

Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang, Min Zhang, Shipei Zeng, Zhongxiang Dai

AI总结:

提出GUI-HARVEST,一种自动工具链优化器,通过证据驱动的诊断和验证,使冻结骨干的GUI智能体自我改进,在多个基准上显著提升性能。

AI中文摘要:

围绕GUI模型的可执行工具链决定了观察如何被组装、动作如何被执行,以及验证、恢复和终止如何被控制。与非GUI智能体的工具链优化相比,自动优化该工具链面临三个相互关联的挑战:协调模型意图与观察到的视觉效果、在可变执行结果下诊断失败,以及识别跨任务的重复失败模式并将其转化为可复用的运行时更改。我们引入了GUI-HARVEST,一种自动工具链优化器,使具有冻结骨干模型的GUI智能体能够自我改进。首先,为了将诊断基于观察到的动作效果,它将模型输出和执行的动作与前后截图对齐,将发现与特定的界面转换联系起来。其次,为了考虑执行的可变性,它将同一任务的重复运行视为一个联合证据单元,使用任务内比较来定位与结果相关的行为差异。第三,它将跨任务的已验证发现整合为重复出现的失败模式,将其映射到有界的源代码编辑,并在评估前记录预测,通过重复执行检查预测的行为效果以及任务性能。在OSWorld-Verified上的实验显示,在六个通用的开放、GUI专用的开放和专有骨干模型上持续获得留出集收益;Qwen3-VL-32B-Instruct在完整套件上提升了12.33个百分点。冻结工具链转移在WindowsAgentArena上以50步将GPT-5提升了13.87个百分点,无需进一步优化。在相同的骨干和初始工具链下,GUI-HARVEST优于Self-Harness和Meta-Harness,表明GUI特定的诊断和验证有助于工具链改进泛化到未见任务。代码可在该https URL获取。

英文摘要:

The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.

补充信息

↑