CUA-Universe:适用于混合GUI+CLI智能体的可扩展动态环境
CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents
- Shanghai Jiao Tong University(上海交通大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出可扩展混合GUI+CLI环境CUA-Universe,通过App-Forge等组件生成数据,训练的9B模型在多基准测试中提升了计算机使用智能体的性能与效率。
AI中文摘要:
计算机使用智能体已在OSWorld、AndroidWorld等基准测试中取得进展,但仍主要通过图形用户界面(GUI)执行操作,常产生低效的操作轨迹。现实中的计算机工作是混合模式,需结合视觉状态检查与精确、高吞吐量的命令行界面(CLI)操作,因此高性能智能体必须在共享应用状态上协调这两种模态。然而,可扩展的混合环境仍较为稀缺,因为为每个真实应用同时支持GUI和CLI通常需要大量手动工程工作。现有智能体也难以互补性地使用两种界面:原生CLI智能体缺乏涉及界面状态或布局任务的视觉感知能力,而原生GUI智能体在适合通过命令执行的操作上效率低下。我们提出CUA-Universe,这是一种可扩展的环境-数据流水线,可将真实桌面软件转化为混合GUI+CLI环境。App-Forge将应用适配为可复现的虚拟机(VM),并对其发现、包装或生成的命令行界面进行处理,可扩展至16个应用;Task-Weave通过对种子文件的可复用操作合成难度可控的多样化混合任务;Path-Steer引导回滚过程沿高效混合路径进行,并收集经验证的轨迹用于后续训练。基于该数据的训练使智能体行为从低效的GUI交互、脆弱的CLI脚本编写转向有效的GUI+CLI协同操作。我们的9B模型在CUA-Verse上提升了成功率与效率(得分+39.3点;步骤数-37%,token数-60%),在OSWorld上(成功率+16.8点;步骤数-57%,token数-44%),在OSWorld-MCP上(得分+7.84点;步骤数-27%,token数-30%)。CUA-Universe为构建更强大、高效的计算机使用智能体提供了可扩展路径。
英文摘要:
Computer-use agents have advanced on benchmarks like OSWorld and AndroidWorld, but still act mostly through the GUI, often producing inefficient trajectories. Real-world computer work is hybrid, combining visual-state inspection with precise, high-throughput command-line operations, so capable agents must coordinate both modalities over shared application state. Yet scalable hybrid environments remain scarce because supporting both GUI and CLI over real applications typically requires substantial manual engineering for each application. Existing agents also struggle to use the two interfaces complementarily: CLI-native agents lack visual perception for tasks involving interface state or layout, while GUI-native agents are inefficient for operations better executed through commands. We introduce CUA-Universe, a scalable environment-to-data pipeline that turns real desktop software into hybrid GUI+CLI environments. App-Forge adapts applications into reproducible VMs and command-line surfaces it discovers, wraps, or generates, scaling to 16 applications; Task-Weave synthesizes diverse hybrid tasks of controllable difficulty from reusable operations over seed files; and Path-Steer steers rollouts along efficient hybrid paths and harvests verified trajectories for post-training. Training on this data shifts behavior from inefficient GUI interaction and brittle CLI scripting toward effective GUI+CLI orchestration. Our 9B model improves both success and efficiency on CUA-Verse (Score +39.3 pts; -37% steps, -60% tokens), OSWorld (SR +16.8 pts; -57% steps, -44% tokens), and OSWorld-MCP (Score +7.84 pts; -27% steps, -30% tokens). CUA-Universe provides a scalable path toward more capable and efficient computer-use agents.