Qwen-UI-Agent技术报告:面向下一代以真实世界为中心的基础图形用户界面智能体
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究推出以真实世界为中心的基础GUI智能体Qwen-UI-Agent,结合多环境与数据飞轮等机制,在移动端基准达最优,在多类任务上表现具竞争力。
AI中文摘要:
图形用户界面(GUI)智能体有潜力成为现有数字设备上的通用执行器。为推动其向真实世界应用迈进,我们设想这类智能体可在真实设备上可靠运行、跨平台执行工作流、将GUI交互与命令行界面(CLI)执行相结合、完成长周期任务、主动启动有用服务,并以最少人力自主提升自身能力。基于这一愿景,我们推出Qwen-UI-Agent,一款覆盖移动端、计算机使用、网页及DeepSearch环境的以真实世界为中心的基础GUI智能体。Qwen-UI-Agent将多样沙箱环境与大规模真实设备移动端运行时相结合,其统一动作空间将GUI操作与CLI执行交织,在单次模型推理中生成批量动作;一种AutoResearch式数据飞轮机制利用智能体构建任务与环境、诊断故障并规划后续迭代;在线强化学习支持对超过100步的轨迹进行训练,超10000个并发环境加速部署;轻量 harness 层支持移动端与计算机端的主动服务启动及有状态工作流。在广泛评估中,Qwen-UI-Agent在移动端使用基准测试中达到当前最优性能,在计算机与浏览器使用任务上与Opus 4.8、Gemini 3.1 Pro、GPT-5.6 Sol等前沿模型相比具备竞争力:移动端使用上,在MobileWorld得82.1%、MobileWorld-Real得92.2%、AndroidDaily得97.5%;计算机使用上,在OSWorld-Verified得79.5%、OSWorld-v2得40.0%部分进度分数;浏览器使用与GUI定位上,在WebArena得73.6%、ScreenSpot-Pro得81.5%。
英文摘要:
GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. Qwen-UI-Agent combines diverse sandbox environments with a large-scale real-device mobile runtime. Its unified action space interleaves GUI operations with CLI execution and generates batched actions in a single model turn. An AutoResearch-style data flywheel uses agents to construct tasks and environments, diagnose failures, and plan subsequent iterations. Online RL supports training on trajectories exceeding 100 turns, with over 10,000 concurrent environments accelerating rollout. A lightweight harness layer supports proactive service initiation and stateful workflows across mobile and computer. Across a broad suite of evaluations, Qwen-UI-Agent sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models, including Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. On mobile use, it achieves 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily. On computer use, it achieves 79.5% on OSWorld-Verified and a 40.0% partial-progress score on OSWorld-v2. On browser use and GUI grounding, it achieves 73.6% on WebArena and 81.5% on ScreenSpot-Pro, respectively.