arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WM-R1:用强化学习训练GUI智能体以推理并利用世界模型

WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning

Yu Han, Tianwen Qian

arXiv 2608.27508首次发表:更新:

发表机构

School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出首个用世界模型替代真实环境训练GUI智能体的强化学习框架WM-R1,在Android基准测试中其性能显著优于仅GRPO的基线方法和推理时模拟方法。

AI 中文摘要

采用强化学习(RL)训练的GUI智能体在移动平台上展现出强大的环境学习能力,但RL通常需要大量真实环境交互,导致资源成本高且不稳定,尤其是在GUI场景中。为解决这些问题,我们提出WM-R1,这是首个用世界模型而非真实环境训练移动GUI智能体的强化学习框架。具体而言,世界模型在所有回合中作为状态转换的来源,替换训练循环内的真实Android环境;WM-R1还将世界模型直接嵌入推理过程,使智能体在确定最终动作前能推理候选动作的后果。关键在于,WM-R1无需真实环境交互,支持基于世界模型的大规模并行化、步级粒度轨迹生成,并引入多维基于规则的奖励,共同优化任务成功率、轨迹效率和世界模型利用率。为实现高效训练,我们整理了包含2000个具有挑战性任务的高质量数据集。在Android移动基准测试上的实验表明,经WM-R1训练的智能体显著优于仅采用GRPO的基线方法和推理时模拟方法。代码可在此https URL获取。

英文摘要

GUI agents trained with reinforcement learning (RL) have showcased strong environment learning capabilities on mobile platforms. However, RL typically demands extensive real-environment interactions, leading to high resource costs and instability, especially in GUI scenarios. To address these, we propose WM-R1, the first reinforcement learning framework that trains mobile GUI agents with world models instead of real environments. Specifically, world models serve as the source of state transitions during all rollouts, replacing the real Android environment within the training loop. WM-R1 also embeds world models directly into the thinking process, enabling agents to reason about the consequences of candidate actions before committing to the final action. Crucially, WM-R1 eliminates the need for real-environment interaction, supports massively parallelized and step-level granularized trajectory generation grounded in world models, and introduces a multi-dimensional rule-based reward that jointly optimizes task success, trajectory efficiency, and world model utilization. For efficient training, we curate a high-quality dataset of 2000 challenging tasks. Experiments on Android mobile benchmarks demonstrate that WM-R1-trained agents significantly outperform GRPO-only baselines and inference-time simulation methods. Code is available at https://github.com/genalyu/WM-R1 .

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑