发表机构
Institute of Information Engineering, Chinese Academy of Sciences; MiLM Plus, Xiaomi Inc.; School of Cyber Security, University of Chinese Academy of Sciences; College of Cryptology and Cyber Science, Nankai University(中国科学院信息工程研究所; 小米公司MiLM Plus; 中国科学院大学网络空间安全学院; 南开大学密码与网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MaP统一框架,通过掩码轨迹预测协调多任务GUI导航,缓解优化冲突与数据异质性,在五个基准上显著优于直接混合训练。
AI 中文摘要
图形用户界面(GUI)智能体自主与软件交互以完成用户请求,其中GUI导航是最关键且最具挑战性的能力。掌握这一能力需要逐步决策、状态-动作对齐和长程规划的复杂协同。虽然直接混合这些对应的导航任务似乎能直观地同时获取这些技能,但这种直接组合严重受限于不一致的优化目标和深度的数据异质性。为克服这些障碍,我们提出MaP(代表“掩码轨迹预测”),一个统一框架,能无缝协调不同的GUI导航任务。通过将多轮GUI交互建模为轨迹,并通过组件掩码和预测定义训练目标,MaP将优化从任务特定的边际分布转变为一致的目标。此外,为处理多个导航任务间的数据异质性,我们设计了一个角色感知适配器学习模块,动态地将每个令牌路由到专门的表示空间。在五个代表性GUI导航基准上的大量实验表明,MaP有效缓解了梯度冲突,显著优于直接混合训练,为多任务GUI导航建立了稳健的范式。
英文摘要
Graphical User Interface (GUI) Agents autonomously interact with software to fulfill user requests, where GUI navigation stands out as the most critical and challenging capability. Mastering this capability demands a complex synergy of step-wise decision-making, state-action alignment, and long-horizon planning. While directly mixing these corresponding navigation tasks seems intuitive to simultaneously acquire these skills, such a direct combination is severely bottlenecked by inconsistent optimization objectives and profound data heterogeneity. To overcome these barriers, we propose the MaP (stands for ``\textbf{M}asked Tr\textbf{a}jectory \textbf{P}rediction''), a unified framework that seamlessly harmonizes divergent GUI navigation tasks. By modeling multi-turn GUI interactions as a trajectory and defining training objectives through component masking and prediction, MaP shifts the optimization from task-specific marginal distributions to a consistent objective. Furthermore, to handle the data heterogeneity across multiple navigation tasks, we design a role-aware adapter learning module that dynamically routes each token to a specialized representation space. Extensive experiments on five representative GUI navigation benchmarks demonstrate that MaP effectively mitigates gradient conflicts and significantly outperforms the direct mixture training, establishing a robust paradigm for multi-task GUI navigation.
CommentsAccepted to EMNLP 2026