发表机构
School of Artificial Intelligence, Shanghai Jiao Tong University; Honor Device Co., Ltd(上海交通大学人工智能学院; 荣耀终端有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对智能体跨设备协同操作难评估的问题,引入DevicesWorld基准测试,整合多类设备环境及众多任务,通过固定评估集评估五个前沿智能体系统,发现成功率低,为可靠跨设备智能体研究提供了可执行、可重现及有诊断作用的评估。
AI 中文摘要
基于语言模型的智能体在操作单个数字环境方面有快速改进,但现实世界用户目标常跨多个设备。现有基准测试多集中于单一执行环境,难以评估智能体跨异构设备获取和整合信息及完成端到端任务的能力。我们引入DevicesWorld,一个用于跨设备协同操作的大规模可执行基准测试。它包含6140个任务,整合了移动、桌面和物联网三类设备环境。我们用固定评估集评估了五个前沿语言模型智能体系统,所有方法成功率都低,最佳仅达12.5%。失败运行中约28.7%至少满足一个评分条件却仍未完成全部任务。DevicesWorld将跨设备协同操作转化为可执行、可重现且对可靠跨设备智能体研究有诊断作用的评估问题。
英文摘要
LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart homes. However, real-world user goals often span multiple devices: information may come from a phone, be processed on a desktop, and the result may need to appear on another device. Most existing benchmarks center on a single dominant execution environment, making it difficult to evaluate whether agents can acquire and integrate information across heterogeneous devices and complete end-to-end tasks with cross-device dependencies. We introduce DevicesWorld, a large-scale executable benchmark for cross-device collaborative operation. DevicesWorld contains 6,140 tasks and integrates three classes of device environments -- mobile, desktop, and IoT -- into a unified cross-device interaction and evaluation framework. Each task defines a natural-language user goal, participating devices and initial states, executable actions, rule-based verifiers, and a cleanup procedure. A multi-stage construction and quality-control pipeline keeps tasks close to realistic user needs while allowing final outcomes to be automatically verified from device states and generated files. We evaluate five frontier LLM-agent systems on a fixed evaluation set. All methods achieve low success rates, with the best reaching only 12.5%. Among failed runs, about 28.7% satisfy at least one scoring condition yet still fail the full task. Trajectories show that agents become stuck acquiring information or manipulating interfaces, confuse source and output devices, or terminate before all conditions are jointly satisfied. DevicesWorld turns cross-device collaborative operation into an executable, reproducible, and diagnostically useful evaluation problem for research on reliable cross-device agents.
Commentshttps://github.com/AgenticOrgLab/DevicesWorld