arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过视觉状态转换扩展 GUI 智能体

Scaling GUI Agents with Visual State Transitions

Xiangyan Liu, Kaixin Li, Haonan Wang, Biao Wu, Meng Fang, Longxu Dou, Chao Du, Michael Qizhe Shieh, Tianyu Pang

arXiv 2607.24112首次发表:更新:

发表机构

NUS; Sea AI Lab; UTS; University of Liverpool(新加坡国立大学; Sea人工智能实验室; 悉尼科技大学; 利物浦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究通过状态转换预训练扩展 GUI 智能体,联合优化逆动力学和正向动力学预训练统一多模态模型,在桌面和移动 GUI 场景基准测试中优于基线,联合动力学优化有稳定改进,下游性能随转换数据量提升。

AI 中文摘要

我们引入状态转换预训练(STP)作为 GUI 智能体的一个新的扩展轴。在 STP 阶段,通过联合优化逆动力学(从状态变化预测动作)和正向动力学(从当前状态和动作预测下一个状态),在视觉状态转换上持续预训练一个统一的多模态模型。这使模型具备更好的基于动作的视觉表示和 GUI 动力学的内部世界模型。在带有任务指令的轨迹上微调时,我们的 STP 训练模型在桌面和移动 GUI 场景的智能体基准测试中始终优于仅通过直接轨迹微调训练的基线。进一步的实证研究表明,联合动力学优化比单目标训练有稳定的改进,并且下游性能随转换数据量稳步扩展。

英文摘要

We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unified multimodal model on visual state transitions by jointly optimizing inverse dynamics (predicting actions from state changes) and forward dynamics (predicting next states from current states and actions). This optimization equips the model with better action-grounded visual representations and an internal world model of GUI dynamics. When subsequently fine-tuned on trajectories with task instructions, our STP-trained models consistently outperform baselines trained solely via direct trajectory fine-tuning across agent benchmarks in both desktop and mobile GUI scenarios (AgentNetBench, AndroidControl, and GUIOdyssey). Further empirical studies show that joint dynamics optimization yields stable improvements over single-objective training, and downstream performance scales steadily with the volume of transition data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑