arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22847cs.AI

GSAR:面向移动GUI智能体的目标-状态-锚点奖励,结合自进化数据合成技术

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

Long Zhang, Yuhan Chen, Chaoran Zhang, Wanxia Cao, Kun Huang, Pengzhi Gao, Wei Liu, Jian Luan, Chenliang Li, Lixin Zou

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出GSAR奖励框架,结合自进化数据合成与状态-锚点机制,解决VLM-GUI智能体训练中数据合成及奖励信号的问题,在多基准测试中表现优异,为GUI智能体训练提供可扩展方案。

中文摘要 AI 辅助

基于视觉语言模型(VLM)的GUI智能体可从在线强化学习(RL)中显著获益,但其训练受两个核心问题瓶颈制约:当前GUI智能体的数据合成方法依赖特定环境,难以生成多样化数据;现有评估器要么可扩展性有限,要么提供的奖励信号不准确、不可靠。为克服这些挑战,我们提出GSAR(Goal-State-Anchor Reward),一种支持可扩展任务生成、能为稳定高效的策略优化提供可靠奖励信号的RL奖励框架。我们的方法具备自进化数据合成能力,可通过任务执行生成多个环境,并产生多样化任务与目标状态;与之互补的状态-锚点机制会自动将成功目标状态中与任务相关的UI元素标注为参考锚点。在RL训练期间,这些参考锚点能提供准确、可扩展的奖励信号,大幅提升训练效率。大量评估表明,我们的框架在离线轨迹验证上准确率超过90%,表现最接近基于规则的方法;此外,使用我们的奖励框架训练的智能体在AndroidWorld及我们构建的基准测试中均展现出优异性能,为GUI智能体训练建立了一种可扩展的方法。

英文摘要

Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental issues: current data synthesis methods for GUI Agents rely on specific environments and struggle to generate diverse data, while existing evaluators either suffer from limited scalability or provide inaccurate and unreliable reward signals. To overcome these challenges, we introduce GSAR (Goal-State-Anchor Reward), a RL reward framework that supports scalable task generation and delivers reliable reward signals for stable and efficient policy optimization. Our approach features self-evolving data synthesis, which produces multiple environments through task execution and generates diverse tasks and goal states. Complementing this, a state-anchor mechanism automatically annotates task-relevant UI elements in successful goal states as reference anchors. During RL training, these reference anchors provide accurate, scalable reward signals that substantially enhance efficiency. Extensive evaluations demonstrate that our framework achieves over 90% accuracy on offline trajectory verification and performs closest to rule-based methods. Furthermore, agents trained using our reward framework exhibit strong performance on both AndroidWorld and our constructed benchmark, establishing a scalable approach for GUI agent training.

发表机构

  • Wuhan University(武汉大学)
  • Xiaomi Inc.(小米公司)

机构由 AI 辅助整理,请以论文原文为准。

↑