发表机构
Beihang University; Beijing Institute of Technology; Baidu Inc.(北京航空航天大学; 北京理工大学; 百度公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
JarvisGUI是一个动态基准测试,通过跨Android、Windows和Ubuntu的多设备工作流评估GUI智能体,揭示现有智能体在状态传递和跨平台推理方面的关键能力差距。
AI 中文摘要
现实世界的图形用户界面(GUI)使用经常涉及跨多个设备和平台的工作流程,需要传递中间结果、维护共享状态,并在异构环境之间进行协调。然而,现有的GUI基准测试绝大多数是在单设备、静态定义的任务上评估智能体,因此这些跨设备能力在很大程度上未被检验,导致对智能体实际应用准备情况的评估过于乐观。我们提出了JarvisGUI,这是一个动态基准测试,用于在需要跨异构平台(包括Android、Windows和Ubuntu)协调交互的跨设备工作流程中评估GUI智能体。具体来说,JarvisGUI将GUI任务表述为轻量级类型系统下的输入-输出转换,这使我们能够自动组合多步骤、跨设备的工作流程,并在统一框架内动态评估智能体性能。通过在跨越多个操作系统的虚拟环境中评估智能体,JarvisGUI揭示了最先进的开源GUI智能体在现实工作流程所需的跨设备状态传递感知、跨平台上下文推理和长周期依赖管理方面存在困难,暴露了现有基准测试无法发现的关键能力差距。
英文摘要
Real-world GUI usage frequently involves workflows that span multiple devices and platforms, requiring the transfer of intermediate results, maintenance of shared state, and coordination across heterogeneous environments. However, existing GUI benchmarks overwhelmingly evaluate agents on single-device, statically defined tasks, thus leaving such cross-device capabilities largely unexamined, resulting in an overly optimistic assessment of agents' readiness for real-world usage. We introduce JarvisGUI, a dynamic benchmark that evaluates GUI agents on cross-device workflows requiring coordinated interaction across heterogeneous platforms, including Android, Windows, and Ubuntu. Specifically, JarvisGUI formulates GUI tasks as input-output transformations under a lightweight type system, which allows us to automatically compose multi-step, cross-device workflows and dynamically evaluate agent performance within a unified framework. By evaluating agents in virtual environments spanning multiple operating systems, JarvisGUI reveals that state-of-the-art open-source GUI agents struggle with the state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management required for real-world workflows, exposing a critical capability gap invisible to existing benchmarks.
CommentsAccepted to EMNLP 2026 (Main Conference)