arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SEE:用于长视野GUI代理轨迹合成的结构感知探索与利用

SEE: Structure-aware Exploring & Exploiting for Long-horizon GUI Agent Trajectory Synthesis

Zhuohang Fan, Beichen Zhang, Yuanfa Li, Changqiao Wu, Wei Liu, Jian Luan, Weigang Zhang

arXiv 2607.18046首次发表:更新:

发表机构

Harbin Institute of Technology, Weihai; Harbin Institute of Technology (Weihai) Qingdao Research Institute(哈尔滨工业大学(威海); 哈尔滨工业大学(威海)青岛研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对GUI代理轨迹合成中缺乏高覆盖率长视野轨迹的问题,提出SEE两阶段数据合成框架,通过高效探索和基于图的合成,产生可重复可解释数据,避免虚假循环,提升代理任务成功率和泛化能力,还将发布代码和数据集。

AI 中文摘要

由视觉语言模型驱动的图形用户界面(GUI)代理有望实现现实世界移动任务的自动化。然而,进展受到从元素丰富且快速发展的应用程序中收集的高覆盖率、长视野交互轨迹缺乏的限制。现有管道通常依赖昂贵的人工演示或策略框架,容易过度采样常见流程,而遗漏罕见转换和复杂多步程序。为解决此问题,我们提出SEE,一个两阶段数据合成框架,包括一个高效探索阶段,在屏幕和元素上构建显式UI转换图,以及一个基于图的合成阶段,通过规划和受控采样组合不同的多步轨迹。这种设计产生可重复和可解释的数据生成,同时明确防止虚假循环并实现长视野组合。在多个真实世界应用程序中,SEE产生平均长度为14.8步的轨迹,同时避免虚假循环,在SEE上微调的代理实现了更高的任务成功率和对未见屏幕的泛化能力。我们将公开发布合成代码和数据集。

英文摘要

Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-coverage, long-horizon interaction trajectories collected from element-rich and rapidly evolving apps. Existing pipelines often rely on costly human demonstrations or on-policy framework, which tends to over-sample common flows while missing rare transitions and complex multi-step procedures. To address this problem, we propose SEE, a two-stage data synthesis framework consisting of (i) an efficient exploration stage that builds an explicit UI transition graph over screens and elements, and (ii) a graph-based synthesis stage that composes diverse multi-step trajectories via planning and controlled sampling. This design yields reproducible and explainable data generation, while explicitly preventing spurious cycles and enabling long-horizon composition. Across multiple real-world apps, SEE produces trajectories with an average length of 14.8 steps while avoiding spurious loops, and agents fine-tuned on SEE achieve improved task success and generalization to unseen screens. We will publicly release our synthesis code and dataset.

CommentsAccepted (Oral) by ACM International Conference on Multimedia 2026 (ACM MM 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑