arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RobotWorld:跨多样任务与具身形态的多模态智能体基准测试

RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments

Zhiqin Yang, Chenxin Li, Xiaomeng Hu, Yibin Liu, Weidong Huang, Jiankai Sun, Haitao Li, Zijian Wu, Yuzhi Huang, Fanding Huang, Hanwen Sun, Jiashun Liu, Jingqi Tong, Mingxin Huang, Shaoli Hu, Shijue Huang, Tianyi Bai, Xinyuan Wang, Yunlong Lin, Zhengyang Tang, Zhexin Zhang, Zhuo Chen, Xierui Song, Juntao Dai, Boyuan Chen, Jiaming Ji, Fangneng Zhan, Mengkang Hu, Wei Xue, Yonggang Zhang, Han Hu, Tsung-Yi Ho, Yike Guo

arXiv 2610.10409首次发表:更新:

AI 中文总结

本文提出RobotWorld仿真测试平台,含84个跨多种机器人的任务,评估多模态智能体物理执行能力,揭示其感知与控制工作流难以稳定组合,并指出不同模型的能力差异与改进目标。

AI 中文摘要

通用智能体日益能够编写代码、使用工具并完成复杂的数字任务,这引发了一个问题:这些能力在多大程度上能够迁移到物理世界中。为了探究这一问题,我们引入了RobotWorld,一个具有挑战性的仿真测试平台,用于机器人使用:将指令和观测通过机器人接口转化为物理任务执行。其84个任务涵盖操作、移动操作、移动、驾驶和空中控制,并带有明确的交互预算和可执行的成功检查。通过将任务结果与执行轨迹相结合进行分析,我们识别出能够迁移的能力以及阻碍可靠完成的差距。此外,我们发现当前智能体能够构建复杂的感知与控制工作流,包括图像分割、相机标定、空间估计和基于动力学的计算。然而,这些能力并不能稳定地组合成成功的行为:智能体尽管到达了指令姿态,却丢失了任务相关的物体状态,未能纠正无效动作,恢复过晚,或将未完成任务误认为完成。这种不均衡的迁移在不同模型间也存在差异:Astra在空间和受限接触目标上更常成功,而Opus 5.5在连续平衡和定时交互目标上更常成功。通过将这些结果与执行行为联系起来,RobotWorld既提供了一个严格的试验场,也提供了对剩余能力差距的经验性说明,从而为训练和设计更可靠的物理世界智能体确立了具体目标。

英文摘要

General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.

Comments62 pages, 25 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑