发表机构
Hunyuan AI Data Team(混元AI数据团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出AutoGUIWorld框架,利用图像生成器与规划器合成GUI交互轨迹,无需运行真实环境,生成7.9万训练样本,微调Qwen3.5-35B-A3B后显著提升OSWorld和ScienceBoard任务性能。
AI 中文摘要
GUI智能体需要高质量的交互轨迹,以学习软件环境如何响应动作、维持状态并支持多步骤工作流。然而,可用轨迹的多样性受到底层环境中可访问的应用程序、界面状态和工作流的限制。扩大这种覆盖范围需要部署日益多样化和复杂的软件,而专门的应用程序会带来额外的安装、配置和运行成本。我们引入了AutoGUIWorld,一个数据生成框架,它结合了图像生成器的视觉先验与规划器的任务知识,无需部署或运行相应的软件环境即可合成GUI交互轨迹。AutoGUIWorld从操作系统上下文、视觉外观和界面状态的结构化规范中采样初始GUI场景,并基于这些场景生成任务。然后,规划器指定原子动作及其预期的视觉结果,而图像生成器迭代地编辑当前截图以生成后续观察。动作落地和转换级质量过滤在Ubuntu、Windows、macOS和Chrome上产生了79,266个带空间标注的步骤级训练样本。在AutoGUIWorld轨迹上微调Qwen3.5-35B-A3B,将OSWorld上的平均任务得分从33.0%提高到40.8%,并将ScienceBoard上的任务成功率从14.0%提高到32.2%。这些结果表明,生成的轨迹提高了GUI智能体在真实桌面和科学任务上的性能。
英文摘要
GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.