发表机构
Jilin University; OPPO(吉林大学; OPPO)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Infinite-Dreamer,利用图像编辑世界模型合成高保真GUI数据,训练Infinite-Actor智能体,在多个基准上显著提升性能。
AI 中文摘要
图形用户界面(GUI)智能体已成为跨多种应用自动化复杂数字工作流的一种有前景的范式。然而,训练高性能且可泛化的智能体从根本上依赖于大规模、高保真的视觉-动作轨迹,而这些轨迹众所周知难以获取。虽然人类演示不可扩展,但现有的GUI世界模型依赖文本描述或HTML渲染,丢弃了关键的像素级视觉细节,如图标和布局样式。为解决此问题,我们引入了Infinite-Dreamer,一种由像素级图像编辑世界模型驱动的免模拟数据合成方法。通过将GUI转换概念化为图像编辑任务,我们利用视觉语言模型(VLMs)将动作引发的UI变化描述为结构化的增量文本。然后,我们微调一个图像编辑骨干网络,以可控地合成逼真的屏幕截图转换。我们利用该模型生成单帧视觉鲁棒性数据和多步想象轨迹。为验证我们方法的有效性,我们仅在合成数据上微调Qwen3-VL基线,获得Infinite-Actor,并在AndroidWorld、MobileWorld和AndroidControl-Curated基准上进行评估。Infinite-Actor在多个规模上持续优于Qwen3-VL基线:Infinite-Actor-8B在AndroidWorld上将Pass@1提高了+4.45,并将MobileWorld的Pass@3成功率几乎翻倍,而Infinite-Actor-2B将Pass@1提高了+9.05。代码可在以下网址获取:https URL。
英文摘要
Graphical User Interface (GUI) agents have emerged as a promising paradigm for automating complex digital workflows across diverse applications. However, training highly capable and generalizable agents fundamentally relies on massive, high-fidelity visual-action trajectories, which are notoriously difficult to acquire. While human demonstrations are unscalable, existing GUI world models rely on text descriptions or HTML rendering, discarding crucial pixel-level visual details like icons and layout styles. To address this issue, we introduce Infinite-Dreamer, a simulation-free data synthesis method powered by a pixel-level Image Editing World Model. By conceptualizing GUI transitions as image editing tasks, we leverage Vision-Language Models (VLMs) to describe action-induced UI changes as structured delta-text. We then fine-tune an image editing backbone to controllably synthesize realistic screenshot transitions. We utilize this model to generate both single-frame visual robustness data and multi-step imaginary trajectories. To validate the effectiveness of our approach, we fine-tune the Qwen3-VL baseline solely on the synthesized data to obtain Infinite-Actor, and evaluate it on AndroidWorld, MobileWorld, and AndroidControl-Curated benchmarks. Infinite-Actor consistently outperforms the Qwen3-VL baselines across scales: Infinite-Actor-8B improves AndroidWorld Pass@1 by +4.45 and nearly doubles the MobileWorld Pass@3 success rate, while Infinite-Actor-2B improves Pass@1 by +9.05. Code is available at https://github.com/swaydy-n/Infinite-Dreamer.