发表机构
Tencent Hy Frontier Team(腾讯 Hy 前沿团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
UI-Mate是集成环境基训练栈与上下文内演示学习的GUI智能体,在多计算机使用基准上创开放权重SOTA,提升了长周期办公任务的执行可靠性
AI 中文摘要
基础图形用户界面(GUI)智能体可自动化复杂数字任务,但其部署受限于稀缺且有偏差的训练数据、模糊的提示以及不可靠的执行。常规工作流程依赖于用户特定工具和默认惯例,因此未明确说明的指令可能会导致运行间出现任意变化。我们提出UI-Mate,这是一种基础GUI智能体,它集成了基于环境的训练栈与上下文内演示学习。UI-Mate有三项贡献:1. 可扩展的基于环境的训练栈:闭环数据引擎通过统一的任务-验证器包,在大规模并行环境中自动执行任务生成、环境构建、 rollout( rollout 指智能体与环境交互生成轨迹的过程)、过滤、能力平衡、监督微调(SFT)以及在线强化学习。2. 上下文内演示学习:一种将多模态演示转换为灵活子任务级工作流的机制,可遵循相关演示步骤并从实时界面重新规划。3. OSWorkerBench基准与见解:该基准包含41个应用中的100项长周期办公任务,支持仅指令和演示引导的评估。其演示资源分为两类:33任务的自演示设置(由相同目标的强智能体成功rollout构建)和45任务的变体演示设置(由相关但非完全相同任务的人类记录构建)。实验表明,UI-Mate-27B在通用计算机使用基准上创下开放权重新SOTA,在OSWorld-Verified上得分为77.0%,在WindowsAgentArena上得分为66.2%;在OSWorkerBench上,其严格成功率达41.0%,进度达76.9%,较其基础模型Qwen3.6-27B分别提升17.7和24.5个百分点;在33任务的自演示子集上,单次演示使严格成功率从17.2%提升至35.4%,进度从67.9%提升至81.1%,大幅提高了长周期任务的可靠性。项目页面:this https URL
英文摘要
Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution. Routine workflows rely on user-specific tools and tacit conventions, so unstated instructions can produce arbitrary variations across runs. We present UI-Mate, a foundation GUI agent that integrates an environment-grounded training stack with in-context demonstration learning. UI-Mate makes three contributions: A Scalable Environment-Grounded Training Stack: A closed-loop data engine automates task generation, environment construction, rollout, filtering, capability balancing, SFT, and online RL across massively parallel environments via unified task-verifier bundles. In-Context Demonstration Learning: A mechanism that transforms multimodal demonstrations into flexible subtask-level workflows, follows relevant demonstrated steps, and re-plans from the live interface. OSWorkerBench Benchmark and Insights: A benchmark of 100 long-horizon office tasks across 41 applications that supports instruction-only and demonstration-guided evaluation. Its demonstration resources separate a 33-task self-demo setting, built from successful strong-agent rollouts of the same targets, from a 45-task variant-demo setting, built from human recordings of related but non-identical tasks. Experiments show that UI-Mate-27B sets a new open-weight state of the art on general computer-use benchmarks, scoring 77.0% on OSWorld-Verified and 66.2% on WindowsAgentArena. On OSWorkerBench, it reaches 41.0% strict success and 76.9% progress, outperforming its Qwen3.6-27B base by 17.7 and 24.5 points. On the 33-task self-demo subset, one demonstration raises strict success from 17.2% to 35.4% and progress from 67.9% to 81.1%, substantially improving long-horizon reliability. Project page: https://ui-mate.github.io.
CommentsUI-Mate Technical Report. Project page: https://ui-mate.github.io