发表机构
Beijing University of Posts and Telecommunications; State Key Laboratory for General Artificial Intelligence, BIGAI; China University of Geosciences (Beijing); Beijing Institute of Technology(北京邮电大学; 通用人工智能国家重点实验室,字节跳动公司人工智能研究院; 中国地质大学(北京); 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究 GUI 智能体长周期任务扩展性差的问题,引入 ParaGUIBench 基准测试及 ParaGUI 智能体,通过实验证明并行执行能提升可分解长周期 GUI 任务的成功率与效率。
AI 中文摘要
图形用户界面(GUI)智能体由大型多模态模型驱动,可感知屏幕状态并通过 GUI 操作执行用户指令。但当前智能体在处理长周期任务时扩展性差。人类会分工并行完成子任务,而 GUI 智能体的并行协调却未受关注。为此引入 ParaGUIBench,这是首个用于多个 GUI 智能体在独立桌面实例上并行执行与协调的基准测试,含多设备 Docker 基础设施、233 个任务的数据集及评估系统。还引入 ParaGUI 智能体,在 ParaGUIBench 上其成功率达 46.4%,优于最强串行基线,且步骤和令牌使用更少,表明并行执行可提高可分解长周期 GUI 任务的成功率和效率。
英文摘要
Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as clicking, typing, and scrolling on desktops and mobile devices. However, current agents scale poorly to long-horizon tasks: actions incur costly LMM inferences, and performance degrades as context grows. Humans divide such workloads among collaborators who complete sub-tasks in parallel. Yet parallel coordination among GUI agents has received little attention. To close this gap, we introduce ParaGUIBench, to our knowledge, the first benchmark dedicated to parallel execution and coordination of multiple GUI agents on separate desktop instances. It consists of three components: a multi-device Docker infrastructure with a shared file system; a dataset of 233 tasks spanning six task categories; and an evaluation system with efficiency metrics, including step reduction ratio and token cost. We further introduce ParaGUI, a planner-worker agent that decomposes GUI tasks and dispatches sub-tasks to concurrent workers on separate desktop instances. On ParaGUIBench, ParaGUI reaches a 46.4% success rate, outperforming the strongest serial baseline (Claude Sonnet 4.6) by 12.9 points while using roughly half the steps and less than half the tokens. These results show that parallel execution can improve both success rate and efficiency on decomposable, long-horizon GUI tasks, pointing to a direction worth further study.
Comments15 pages, 5 figures. Project page: https://github.com/pkgunboat/ParaGUIBench