arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02352cs.LGcs.CL

Qwen-CUA:面向几乎所有场景的原生计算机使用智能体

Qwen-CUA: Native Computer Use for (almost) Everything

  • Qwen Team(通义千问团队)
  • Xlang Lab(Xlang实验室)

机构由 AI 辅助整理,请以论文原文为准。

Dunjie Lu, Shuai Bai, Tianyi Bai, Sicheng Fan, Chang Gao, Jian Guan, Feng Hu, Mianqiu Huang, Xingyang Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Ning Li, Day… 展开作者

Dunjie Lu, Shuai Bai, Tianyi Bai, Sicheng Fan, Chang Gao, Jian Guan, Feng Hu, Mianqiu Huang, Xingyang Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Ning Li, Dayiheng Liu, Shixuan Liu, Zheng Liu, Que Shen, Bowen Wang, Junli Wang, Chencan Wu, Rui Xie, Tianbao Xie, Zhihui Xie, Haiyang Xu, An Yang, Tao Yu, Wenzhen Yuan, Xi Zhang, Zhenru Zhang, Mingkang Zhu, Zhaoqing Zhu, Yizhong Cao, Kai Dang, Binyuan Hui, Kaixin Li, Junyang Lin, Haiquan Wang, Zekun Wang, Yiheng Xu, Fan Yan, Mengqi Yuan, Danyang Zhang, Jiajun Zhang, Zhipeng Zhang, Fan Zhou, Fan Zhou

AI总结:

本文提出原生计算机使用智能体Qwen-CUA,采用397B-A17B Qwen混合专家主干,经大规模训练后在多基准测试中性能优于Qwen3.7,为通用能力智能体奠定基础。

AI中文摘要:

原生计算机使用为智能体提供了通用接口,可操作几乎所有人类可用的软件,但需要长程状态跟踪、大规模交互经验,以及从稀疏但可验证的结果中学习。本文介绍Qwen-CUA,这是一个原生计算机使用智能体,采用397B-A17B规模的Qwen混合专家主干模型。它仅通过屏幕截图进行观测,通过键盘和鼠标事件执行动作,无需DOM树、可访问性元数据或任务特定API。其脚手架最多维护20张活跃截图,将较早的视觉历史折叠为固定大小的块,以保留近期证据同时保留可复用的提示前缀。训练阶段,我们构建了可访问近100,000个vCPU和数万个并发环境的云部署集群,构建了约40,000个可验证任务,并收集了日常及专业软件中的个性化长程工作流。我们通过可验证奖励和轨迹切片优化完整轨迹,迭代训练过程会更新监督数据并重新校准强化学习任务。在8个基准测试中,Qwen-CUA的性能优于Qwen3.7,且与领先的专有系统具有竞争力,在OSWorld-Verified上达到86.2,在OSWorld 2.0上的二元完成得分为18.5、部分完成为48.4。将相同方案扩展至参数规模超1万亿的模型得到Qwen-CUA-Max,这些得分分别提升至87.6、21.2和53.3。相较于Qwen3.7,Qwen-CUA还将RedTeamCUA攻击成功率从36.6降至16.4。效率分析、浏览器部署及Bash增强实验进一步刻画了其实际行为。这些结果确立了原生计算机使用是通用能力智能体的基础,并强调可扩展的可验证交互与混合工具使用是关键方向。

英文摘要:

Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes. We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone. It observes only screenshots and acts through keyboard and mouse events, without DOM trees, accessibility metadata, or task-specific APIs. Its scaffold maintains up to 20 active screenshots and folds older visual history in fixed-size blocks to retain recent evidence while preserving reusable prompt prefixes. For training, we build a cloud rollout fleet with access to nearly 100,000 vCPUs and tens of thousands of concurrent environments, construct approximately 40,000 verifiable tasks, and collect personalized long-horizon workflows across everyday and professional software. We optimize complete trajectories with verifiable rewards and trajectory slicing, while iterative training runs refresh supervised data and recalibrate reinforcement-learning tasks. Across eight benchmarks, Qwen-CUA outperforms Qwen3.7 and remains competitive with leading proprietary systems, reaching 86.2 on OSWorld-Verified and 18.5/48.4 binary/partial completion on OSWorld 2.0. Scaling the same recipe to a model with over one trillion parameters yields Qwen-CUA-Max, improving these scores to 87.6 and 21.2/53.3. Qwen-CUA also reduces RedTeamCUA attack success from 36.6 to 16.4 relative to Qwen3.7. Efficiency analyses, a browser deployment, and Bash-augmented experiments further characterize practical behavior. These results establish native computer use as a broadly capable agent foundation and highlight scalable verifiable interaction and hybrid tool use as key directions.

补充信息

↑