arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.06556cs.RO

机器人需要的不仅仅是VLA和世界模型

Robots Need More than VLA and World Models

Elis Karcini, Faisal Mehrban, Quang Nguyen, Mac Schwager, Arash Ajoudani, Cesar Cadena, Jan Peters, Marco Hutter, Haitham Bou-Ammar

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文认为机器人通用智能的关键瓶颈不仅是策略学习,还缺乏将非结构化行为数据转化为机器人可用监督的机制,并提出了四种缺失的接口组件。

中文摘要 AI 辅助

通用机器人智能通常被框定为策略扩展问题:收集更多机器人演示,训练更大的视觉-语言-动作(VLA)模型,并期望更广泛的泛化。在这篇立场论文中,我们认为这种框架是不完整的。核心瓶颈不仅是策略学习,而是缺乏将世界上丰富的非结构化行为数据转化为有监督的机器人监督的机制。人类运动、互联网视频、仿真 rollout 和交互式演示包含关于任务、目标、接触、失败和物理约束的丰富信息,然而这些信息中的大部分无法直接被机器人策略使用,因为它们缺乏特定于具身的动作标签、任务语义和奖励结构。我们为下一代机器人识别了四个缺失的组件:用于自动标注非结构化行为的数据接口、用于将人类运动重定向到机器人动作的具身接口、用于物理接地3D推理的世界模型接口,以及用于从视频和语言推断任务进展和成功的奖励接口。我们调查了机器人基础模型、跨具身数据集、从视频学习、世界模型和奖励建模方面的最新进展,并提出了一个研究议程,以构建不仅能够从机器人演示中学习,而且能够从更广泛的物理世界中学习的机器人系统。

英文摘要

Generalist robot intelligence is often framed as a policy-scaling problem: collect more robot demonstrations, train larger Vision-Language-Action (VLA) models, and expect broader generalisation. In this position paper, we argue that this framing is incomplete. The central bottleneck is not only policy learning, but the absence of mechanisms that convert the world's abundant unstructured behavioural data into grounded robot supervision. Human motion, internet video, simulation rollouts, and interactive demonstrations contain rich information about tasks, goals, contacts, failures, and physical constraints, yet most of this information is not directly usable by robot policies because it lacks embodiment-specific action labels, task semantics, and reward structure. We identify four missing components for the next generation of robotics: data interfaces for autolabelling unstructured behaviour, embodiment interfaces for retargeting human motion to robot actions, world-model interfaces for physics-grounded 3D reasoning, and reward interfaces for inferring task progress and success from video and language. We survey recent progress in robot foundation models, cross-embodiment datasets, learning from video, world models, and reward modelling, and propose a research agenda for building robotics systems that can learn not only from robot demonstrations, but from the broader physical world.

发表机构

  • Motoniq.ai
  • Stanford University(斯坦福大学)
  • Istituto Italiano di Tecnologia(意大利技术研究院)
  • ETH Zurich(苏黎世联邦理工学院)
  • Technical University of Darmstadt(德累斯顿技术大学)
  • UCL Centre for AI(伦敦大学学院人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

↑