arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HyMobileAgent:用于高效GUI代理的数据-环境协同扩展

HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents

Hy Vision Team, Huawen Shen, Zhengyang Tang, Shangpin Peng, Liang Wu, Anran Zhang, Weinong Wang, Yiduo Guo, Chenxin Li, Zhengyao Fang, Yang Ding, Junyi Li, Fei Tang, Zheng Ruan, Yi Zhang, Xingran Zhou, Dingchen Yang, Sunqi Fan, Zhiyi Wan, Han Hu, Xin Lai, Pengyuan Lyu, Chengquan Zhang

arXiv 2607.14548首次发表:更新:

发表机构

PhoneWorld Mock App Factory(手机世界模拟应用工厂)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究移动GUI代理在数字环境中的运行,基于Hy3.0-VL-A3B构建HyMobileAgent,开发数据和环境协同扩展框架,集成多种机制和管道,提出渐进式训练方法,解决移动交互关键瓶颈。

AI 中文摘要

随着大型多模态模型从理解内容转向在数字环境中运行,移动GUI成为数字具身智能的一个具有挑战性且重要的测试平台。移动代理在三个耦合约束下运行。本报告提出了HyMobileAgent,它基于Hy3.0-VL-A3B构建。我们开发了一个以数据和环境为中心的联合扩展框架来解决移动交互的关键瓶颈。该框架集成了GUI感知飞轮、知识管道等。还介绍了由中期训练、监督微调等组成的渐进式训练方法。

英文摘要

As large multimodal models move from understanding content to operating on digital environments, mobile GUI has emerged as a challenging and consequential testbed for digital embodied intelligence. Mobile agents operate under three coupled constraints: precise perception of complex interfaces, scalable acquisition of high-quality interaction data, and robust long-horizon decision making under compounding execution errors. This report presents HyMobileAgent, a mobile GUI agent built on Hy3.0-VL-A3B, a vision-native foundation model featuring native any-resolution input, an A3B-scale deployment budget, and a 32K context window to model extended interaction histories. Rather than relying solely on model scaling, we develop a joint data and environment centric scaling framework to address the key bottlenecks of mobile interaction. Our framework integrates a GUI perception flywheel combining mock-interface synthesis, rejection sampling, and icon-specific augmentation; a knowledge pipeline that transforms tutorial videos into structured interaction data; a million-scale action data pipeline deployed across more than 2000 sandbox and real-device instances with automated failure attribution; the PhoneWorld Mock App Factory, providing a resettable training environment with 34 mock applications and over 34000 tasks; and a structured Planning-and-Reflection mechanism with explicit dead-loop detection for reliable long-horizon execution. We also introduce a progressive training recipe consisting of mid-training, supervised fine-tuning, and reinforcement learning with task-specific reward designs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑