发表机构
Institute of Information Engineering, Chinese Academy of Sciences; Tencent; School of Cyber Security, University of Chinese Academy of Sciences; VCIP & TMCC & DISSec, College of Computer Science, Nankai University(中国科学院信息工程研究所; 腾讯; 中国科学院大学网络空间安全学院; 南开大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GUI-SD-v2,通过两阶段训练框架将在线策略自蒸馏从GUI定位扩展到多轮交互,解决特权遵循和指导不足问题,在AndroidWorld和MobileWorld基准上显著提升Pass@1和Pass@3成功率。
AI 中文摘要
图形用户界面(GUI)智能体通过与软件环境的多轮交互来完成复杂的用户指令,这需要逐步推理和长时记忆来分别指导动作并保留任务相关信息。最近的在线策略自蒸馏(OPSD)方法在GUI定位(GUI智能体的基础子任务)上取得了强劲性能,这得益于来自特权条件自教师模型的密集词级监督。然而,将现有OPSD方法扩展到多轮GUI智能体受到自教师模型有限的特权遵循能力和不足的特权指导的阻碍。在本文中,我们介绍了GUI-SD-v2,即GUI-SD的下一代版本,它将OPSD从GUI定位扩展到多轮GUI交互,并通过两阶段训练框架解决了关键限制。具体而言,GUI-SD-v2首先通过联合优化来自相同GUI状态的有无特权指导的轨迹来加强特权遵循。此外,它通过特权条件自教师模型选择性地蒸馏逐步推理和记忆指导,以支持动作决策和保留后续交互所需的任务相关信息。在两个代表性GUI智能体基准测试AndroidWorld和MobileWorld上进行的大量实验表明,GUI-SD-v2与现有OPSD基线相比表现优异,同时在Pass@1和Pass@3成功率上持续优于所评估的最先进方法。代码和训练数据将公开发布。
英文摘要
Graphical User Interface (GUI) agents enable the fulfillment of complex user instructions through multi-turn interactions with software environments, requiring step-wise reasoning and long-horizon memory to guide actions and retain task-relevant information, respectively. Recent on-policy self-distillation (OPSD) methods have achieved strong performance on GUI grounding, a foundational subtask for GUI agents, owing to dense token-level supervision from privilege-conditioned self-teachers. However, extending existing OPSD methods to multi-turn GUI agents is hindered by self-teachers' limited privilege-following ability and insufficient privileged guidance. In this paper, we introduce GUI-SD-v2, the next version of GUI-SD, which extends OPSD from GUI grounding to multi-turn GUI interaction and addresses key limitations through a two-stage training framework. Specifically, GUI-SD-v2 first strengthens privilege following by jointly optimizing rollouts with and without privileged guidance from the same GUI states. Furthermore, it selectively distills step-specific reasoning and memory guidance through a privilege-conditioned self-teacher, supporting action decisions and the retention of task-relevant information for subsequent interactions. Extensive experiments on two representative GUI agent benchmarks, AndroidWorld and MobileWorld, show that GUI-SD-v2 compares favorably with existing OPSD baselines while consistently outperforming the evaluated state-of-the-art methods in both Pass@1 and Pass@3 success rates. Code and training data will be publicly released.
CommentsUnder Review