arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向全维度GUI智能体导航:基于掩码轨迹预测

Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction

Yan Zhang, Pei Fu, Daiqing Wu, Huawen Shen, Ruoceng Zhang, Shaojie Zhang, Jiahui Yang, Yu Zhou, Can Ma, Zhenbo Luo, Jian Luan

arXiv 2609.25769首次发表:更新:

发表机构

Institute of Information Engineering, Chinese Academy of Sciences; MiLM Plus, Xiaomi Inc.; School of Cyber Security, University of Chinese Academy of Sciences; College of Cryptology and Cyber Science, Nankai University(中国科学院信息工程研究所; 小米公司MiLM Plus; 中国科学院大学网络空间安全学院; 南开大学密码与网络空间安全学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MaP统一框架,通过掩码轨迹预测协调多任务GUI导航,缓解优化冲突与数据异质性,在五个基准上显著优于直接混合训练。

AI 中文摘要

图形用户界面(GUI)智能体自主与软件交互以完成用户请求,其中GUI导航是最关键且最具挑战性的能力。掌握这一能力需要逐步决策、状态-动作对齐和长程规划的复杂协同。虽然直接混合这些对应的导航任务似乎能直观地同时获取这些技能,但这种直接组合严重受限于不一致的优化目标和深度的数据异质性。为克服这些障碍,我们提出MaP(代表“掩码轨迹预测”),一个统一框架,能无缝协调不同的GUI导航任务。通过将多轮GUI交互建模为轨迹,并通过组件掩码和预测定义训练目标,MaP将优化从任务特定的边际分布转变为一致的目标。此外,为处理多个导航任务间的数据异质性,我们设计了一个角色感知适配器学习模块,动态地将每个令牌路由到专门的表示空间。在五个代表性GUI导航基准上的大量实验表明,MaP有效缓解了梯度冲突,显著优于直接混合训练,为多任务GUI导航建立了稳健的范式。

英文摘要

Graphical User Interface (GUI) Agents autonomously interact with software to fulfill user requests, where GUI navigation stands out as the most critical and challenging capability. Mastering this capability demands a complex synergy of step-wise decision-making, state-action alignment, and long-horizon planning. While directly mixing these corresponding navigation tasks seems intuitive to simultaneously acquire these skills, such a direct combination is severely bottlenecked by inconsistent optimization objectives and profound data heterogeneity. To overcome these barriers, we propose the MaP (stands for ``\textbf{M}asked Tr\textbf{a}jectory \textbf{P}rediction''), a unified framework that seamlessly harmonizes divergent GUI navigation tasks. By modeling multi-turn GUI interactions as a trajectory and defining training objectives through component masking and prediction, MaP shifts the optimization from task-specific marginal distributions to a consistent objective. Furthermore, to handle the data heterogeneity across multiple navigation tasks, we design a role-aware adapter learning module that dynamically routes each token to a specialized representation space. Extensive experiments on five representative GUI navigation benchmarks demonstrate that MaP effectively mitigates gradient conflicts and significantly outperforms the direct mixture training, establishing a robust paradigm for multi-task GUI navigation.

CommentsAccepted to EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑