发表机构
ETH Zurich; IBM Research Zurich; Microsoft(苏黎世联邦理工学院; IBM苏黎世研究院; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DeskForge 通过可控组合真实桌面环境生成大规模密集监督,构建 DeskForge-1M 数据集,微调四个视觉语言模型,显著提升 GUI 定位和长时程任务完成性能。
AI 中文摘要
计算机使用智能体需要在复杂的桌面场景中可靠地定位动作目标,在这些场景中,多个应用程序、重叠窗口和视觉上相似的控件会竞争注意力。现有的训练数据很少将此类场景与密集标注配对,或以受控方式对其进行变化。我们引入了 DeskForge,一个可控的桌面环境,它组合并探索真实应用程序,为计算机使用智能体生成大规模监督。它变化应用程序状态、内容、窗口布局、外观和分辨率,并将截图、无障碍树和窗口几何信息融合为密集元素标注,同时记录每个执行动作的结果。利用该环境,我们构建了 DeskForge-1M,一个包含 120 万条带标注的桌面观测数据的语料库,其中包含 1.597 亿个元素实例。我们在从 DeskForge-1M 中抽取的 20 万个定位示例上微调了四个视觉语言模型。所有四个模型在保留的桌面条件和所有五个外部 GUI 定位基准上均有所提升;对于 Qwen3.5-4B,在 ScreenSpot-Pro 上准确率提高了 11.51 个百分点,在 OSWorld-G 上提高了 10.11 个百分点。这些提升也转化为长时程任务完成:在固定规划器下,微调后的动作模型解决了更多 WebArena-Infinity 和 OpenApps 任务,Qwen3.5-4B 分别从 119 个任务中的 31 个增加到 50 个,以及从 100 个任务中的 3 个增加到 15 个。这些结果表明,可控组合真实桌面环境为改进 GUI 定位和长时程计算机使用提供了可扩展的监督来源。框架代码、数据集和微调模型可从项目页面获取:此 https URL
英文摘要
Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. Using this environment, we construct DeskForge-1M, a corpus of 1.2M annotated desktop observations containing 159.7M element instances. We fine-tune four vision-language models on 200K grounding examples drawn from DeskForge-1M. All four improve across held-out desktop conditions and on all five external GUI grounding benchmarks; for Qwen3.5-4B, accuracy increases by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. The gains also translate to long-horizon task completion: under a fixed planner, the fine-tuned action models solve more WebArena-Infinity and OpenApps tasks, with Qwen3.5-4B increasing from 31 to 50 of 119 tasks and from 3 to 15 of 100 tasks, respectively. These results show that controllable composition of real desktop environments provides a scalable source of supervision for improving both GUI grounding and long-horizon computer use. The framework code, the dataset, and the fine-tuned model are available from the project page: https://saidgurbuz.github.io/deskforge/
Comments37 pages, 15 figures, 12 tables. Project page: https://saidgurbuz.github.io/deskforge/