UI-Venus-2 技术报告
UI-Venus-2 Technical Report
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
UI-Venus-2是一款通用型基础GUI智能体,通过统一闭环框架覆盖多环境,在环境、任务、验证维度扩展规模并加入安全机制,为实际应用智能体发展提供开源基础。
AI中文摘要:
多模态图形用户界面(GUI)智能体已成为数字任务自动化的有前景范式,但从面向基准的模型过渡到可靠的实际应用仍面临挑战,原因在于环境覆盖有限、任务构建脆弱以及奖励验证不可靠。本研究提出UI-Venus-2,这是一款通用型基础GUI智能体,旨在通过统一的闭环推理-行动框架在移动、网页和桌面环境中运行。为弥合实际部署的差距,我们在三个关键维度同步扩展:(1)环境:将覆盖范围扩大至170多个多语言移动应用和原生桌面操作系统;(2)任务:采用深度研究流程生成基于功能的指令;(3)验证:采用带有视觉关键点和多模型投票的轨迹级与样本级评估器,为训练提供可靠的强化学习信号。此外,我们整合了安全感知机制,以确保关键行动的受控执行。通过提供一款强大、高效且开源的基础模型,UI-Venus-2推动该领域向更具泛化性、可验证性和自我反思能力的实际应用智能体发展。
英文摘要:
Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.