Qwen-CUA:面向几乎所有场景的原生计算机使用智能体
Qwen-CUA: Native Computer Use for (almost) Everything
- Qwen Team(通义千问团队)
- Xlang Lab(Xlang实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出原生计算机使用智能体Qwen-CUA,采用397B-A17B Qwen混合专家主干,经大规模训练后在多基准测试中性能优于Qwen3.7,为通用能力智能体奠定基础。
AI中文摘要:
原生计算机使用为智能体提供了通用接口,可操作几乎所有人类可用的软件,但需要长程状态跟踪、大规模交互经验,以及从稀疏但可验证的结果中学习。本文介绍Qwen-CUA,这是一个原生计算机使用智能体,采用397B-A17B规模的Qwen混合专家主干模型。它仅通过屏幕截图进行观测,通过键盘和鼠标事件执行动作,无需DOM树、可访问性元数据或任务特定API。其脚手架最多维护20张活跃截图,将较早的视觉历史折叠为固定大小的块,以保留近期证据同时保留可复用的提示前缀。训练阶段,我们构建了可访问近100,000个vCPU和数万个并发环境的云部署集群,构建了约40,000个可验证任务,并收集了日常及专业软件中的个性化长程工作流。我们通过可验证奖励和轨迹切片优化完整轨迹,迭代训练过程会更新监督数据并重新校准强化学习任务。在8个基准测试中,Qwen-CUA的性能优于Qwen3.7,且与领先的专有系统具有竞争力,在OSWorld-Verified上达到86.2,在OSWorld 2.0上的二元完成得分为18.5、部分完成为48.4。将相同方案扩展至参数规模超1万亿的模型得到Qwen-CUA-Max,这些得分分别提升至87.6、21.2和53.3。相较于Qwen3.7,Qwen-CUA还将RedTeamCUA攻击成功率从36.6降至16.4。效率分析、浏览器部署及Bash增强实验进一步刻画了其实际行为。这些结果确立了原生计算机使用是通用能力智能体的基础,并强调可扩展的可验证交互与混合工具使用是关键方向。
英文摘要:
Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes. We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone. It observes only screenshots and acts through keyboard and mouse events, without DOM trees, accessibility metadata, or task-specific APIs. Its scaffold maintains up to 20 active screenshots and folds older visual history in fixed-size blocks to retain recent evidence while preserving reusable prompt prefixes. For training, we build a cloud rollout fleet with access to nearly 100,000 vCPUs and tens of thousands of concurrent environments, construct approximately 40,000 verifiable tasks, and collect personalized long-horizon workflows across everyday and professional software. We optimize complete trajectories with verifiable rewards and trajectory slicing, while iterative training runs refresh supervised data and recalibrate reinforcement-learning tasks. Across eight benchmarks, Qwen-CUA outperforms Qwen3.7 and remains competitive with leading proprietary systems, reaching 86.2 on OSWorld-Verified and 18.5/48.4 binary/partial completion on OSWorld 2.0. Scaling the same recipe to a model with over one trillion parameters yields Qwen-CUA-Max, improving these scores to 87.6 and 21.2/53.3. Qwen-CUA also reduces RedTeamCUA attack success from 36.6 to 16.4 relative to Qwen3.7. Efficiency analyses, a browser deployment, and Bash-augmented experiments further characterize practical behavior. These results establish native computer use as a broadly capable agent foundation and highlight scalable verifiable interaction and hybrid tool use as key directions.