arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23863cs.RO

接地动作模型:3D接地作为机器人学的基础

Grounded Action Model: 3D Grounding as a Foundation for Robotics

Gehao Zhang, Weikai Huang, Shailesh Shailesh, Yiyan Peng, Jiafei Duan, Ranjay Krishna

首次发表
浏览论文内容

中文总结 AI 辅助

提出接地动作模型(GAMs),通过3D接地将语言、点或框提示转换为对象中心表示,预测动作块,在多个基准和真实机器人上显著提升成功率,可作低级控制器支持长时域操作。

中文摘要 AI 辅助

操作策略必须知道哪些物体重要以及它们在哪里,然而当前机器人基础模型所依赖的预训练骨干网络,从视觉-语言-动作模型(VLAs)中的语言到世界-动作模型(WAMs)中的视频生成,并不直接要求这种度量接地,而是留给从机器人演示中隐式学习。我们提出了接地动作模型(GAMs),一种基于3D接地构建的机器人基础模型新范式。GAM可以通过语言、点或框提示进行条件化,这些提示首先被转换为所选对象的共享对象中心表示。该表示捕获目标聚焦的视觉特征和度量对象几何,通过多流变压器与机器人状态历史混合以预测动作块。尽管GAM可以自主运行,它们也可以作为低级控制器,由高级规划器使用其各种输入模态进行控制,从而允许长时域和依赖记忆的操作。在RoboTwin 2.0上,GAM在50个任务中实现了55.3%的平均成功率(而Spatial Forcing为52.0%),在场景随机化下为47.6%(而Abot-M0为30.4%),其动作策略仅在干净场景演示上训练。在LIBERO-PRO上,它在16种扰动设置中实现了最先进的61%平均成功率(而$\pi_{0.5}$为53%),在目标被重新定位或新指定时增益最大。在两台真实机器人上,GAM在双臂YAM上的视觉偏移下保持了17/20的成功率,而$\pi_{0.5}$为4/20,而与Franka上的Molmo2规划器组合在长时域和依赖记忆任务上实现了64.7%的ID和49.8%的OOD步骤完成率。

英文摘要

Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned using language, points, or box prompts, which are first transformed into a shared object-centric representation of the selected objects. This representation captures target-focused visual features and metric object geometry, which is mixed with robot state history through a multi-stream transformer to predict action chunks. Although GAMs can be run autonomously, they can also serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation. On RoboTwin 2.0, GAM achieves an average success rate of 55.3% across 50 tasks (vs. 52.0% for Spatial Forcing), including 47.6% under scene randomization (vs. 30.4% for Abot-M0), with its action policy trained only on clean-scene demonstrations. On LIBERO-PRO, it achieves a state-of-the-art average success rate of 61% (vs. 53% for $π_{0.5}$) across 16 perturbation settings, with the largest gains when targets are relocated or newly designated. On two real robots, GAM retains 17/20 successes under visual shift on a bimanual YAM versus 4/20 for $π_{0.5}$, while its composition with a Molmo2 planner on a Franka achieves 64.7% ID and 49.8% OOD step completion on long-horizon and memory-dependent tasks.

发表机构

  • Northwestern University(西北大学)
  • University of Washington(华盛顿大学)
  • National University of Singapore(新加坡国立大学)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑