arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BlueLM-GUI 技术报告:面向真实设备的自改进移动 GUI 智能体飞轮

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

Tong Ye, Kunyang Han, Guozhi Wang, Longqiang Luo, Zhifeng Ding, Yongxiang Zhang, Xiaolei Shen, Yuxuan Zhang, Zhuping Zhang, Tao Xu, Yue Pan, Yucheng Zhao, Yupei Hu, Yuanjiang Ouyang, Danfeng Shen, Runqi Lin, Hongda Cai, Zhaoxiong Wang, Mengjia Yan, Yingjie Zhong, Chen Zhou, Zeyu Zhang, Xuwen Zhu, Penggang Shi, Mingcheng Luo, Ziyang Wu, Min Jin, Mingfu Shen, Zairong Xu, Fan Zhang, Hao Wang, Liang Liu, Zhulin Xie, Lijun Yao, Xiao Liang, Liangmin Wen, Liqiang Feng, Feilong Wu, Min Hu, Min Chen, Guanjing Xiong, Xiaohu Ruan, Xiaoxin Chen

arXiv 2609.12394首次发表:更新:

发表机构

vivo AI Lab(维沃人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对移动GUI智能体部署中的分布失配、故障利用不足和基准饱和问题,提出以真实设备为中心的飞轮方法,通过双轨数据、三阶段训练和动态基准实现自我改进,在MobileGUI-VBench和AndroidWorld上取得领先成绩。

AI 中文摘要

移动 GUI 智能体正从多模块框架转向端到端训练的原生模型,然而工业部署仍面临三个持续存在的差距。沙盒训练导致与生产环境的分布不匹配;昂贵的真实设备故障仍未得到充分利用;固定基准趋于饱和,失去了指导迭代的能力。我们提出 BlueLM-GUI,一个 35B-A3B 规模的移动 GUI 智能体,构建为以真实设备为中心的飞轮,通过三项原则弥合这些差距。每个样本都重要:采用双轨流水线,包含异构三重系统共识评估和错误纠正与衍生模块,将每条轨迹转化为可用的监督信号。每次部署都是真实的:采用三阶段方案——持续预训练、监督微调和在数百台真实手机上的智能体强化学习——使每次部署都扎根于真实生产环境,从而使模型学到的能力直接迁移到部署中。每个查询都在演进:采用基于配额的三正交轴基准方法,实现精确归因,并允许基准随模型改进而系统升级。BlueLM-GUI 在 MobileGUI-VBench 上达到 87.4,超越最佳闭源模型 5.1 分,在 AndroidWorld 上达到 84.9,是开源模型中的最佳结果,并与闭源模型竞争。这些结果表明,将模型训练和迭代改进扎根于真实设备和三项 Every 原则,能够产生强大、稳健且可迁移的移动 GUI 能力。

英文摘要

Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losing the power to guide iteration. We present BlueLM-GUI, a 35B-A3B mobile GUI agent built as a real-device-centric flywheel that closes these gaps through three principles. Every Sample Matters: a dual-track pipeline with Heterogeneous Triple-System Consensus evaluation and an Error Correction \& Derivation Module salvages every trajectory into usable supervision. Every Rollout Is Real: a three-stage recipe---continual pre-training, supervised fine-tuning, and agentic reinforcement learning on hundreds of real phones---grounds every rollout in real production environments, so the capability the model learns transfers directly to deployment. Every Query Evolves: a quota-driven benchmark methodology with three orthogonal axes enables precise attribution and allows the benchmark to be systematically upgraded as the model improves. BlueLM-GUI achieves 87.4 on MobileGUI-VBench, surpassing the best closed-source model by 5.1 points, and 84.9 on AndroidWorld, the best result among open-source models and competitive with closed-source models. These results demonstrate that grounding model training and iterative improvement in both real devices and the three Every principles yields strong, robust, and transferable mobile GUI capability.

Comments49 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑