arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00028cs.AIcs.CLcs.CVcs.LG

UI-Venus-2 技术报告

UI-Venus-2 Technical Report

  • Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhe… 展开作者

Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhen Lin, Zhenyu Ma, Changhua Meng, Han Peng, Xinyu Qiu, Shuheng Shen, Zhongyi Shui, Weiqiang Wang, Ming Wen, Zhuoer Xu, Hang Yan, Kaiwen Yang, Ruilin Yao, Nanjun Yu, Zhengwen Zeng, Lianrui Zhang, Yunzhu Zhang, Zhe Zhao, Beitong Zhou

AI总结:

UI-Venus-2是一款通用型基础GUI智能体,通过统一闭环框架覆盖多环境,在环境、任务、验证维度扩展规模并加入安全机制,为实际应用智能体发展提供开源基础。

AI中文摘要:

多模态图形用户界面(GUI)智能体已成为数字任务自动化的有前景范式,但从面向基准的模型过渡到可靠的实际应用仍面临挑战,原因在于环境覆盖有限、任务构建脆弱以及奖励验证不可靠。本研究提出UI-Venus-2,这是一款通用型基础GUI智能体,旨在通过统一的闭环推理-行动框架在移动、网页和桌面环境中运行。为弥合实际部署的差距,我们在三个关键维度同步扩展:(1)环境:将覆盖范围扩大至170多个多语言移动应用和原生桌面操作系统;(2)任务:采用深度研究流程生成基于功能的指令;(3)验证:采用带有视觉关键点和多模型投票的轨迹级与样本级评估器,为训练提供可靠的强化学习信号。此外,我们整合了安全感知机制,以确保关键行动的受控执行。通过提供一款强大、高效且开源的基础模型,UI-Venus-2推动该领域向更具泛化性、可验证性和自我反思能力的实际应用智能体发展。

英文摘要:

Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.

↑