发表机构
University of Hong Kong; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; Huazhong University of Science and Technology; Beijing University of Aeronautics and Astronautics; Infiforce(香港大学; 中国科学院深圳先进技术研究院; 华中科技大学; 北京航空航天大学; 英飞拓)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言-动作模型缺乏目标对象细粒度3D理解的问题,提出以对象为中心的3D表示对齐框架,利用SAM3D提供3D先验,经定位、生成掩码、提取表示等步骤与视觉特征对齐,实验证明其可提升模型性能,尤其在长时操纵场景有效。
AI 中文摘要
视觉-语言-动作(VLA)模型在通用机器人操纵方面显示出强大潜力,但现有模型大多依赖2D视觉语言主干,缺乏对目标对象的细粒度3D理解。我们提出了一个以对象为中心的3D表示对齐框架,基于π0构建,使用SAM3D作为冻结的3D教师在训练期间提供目标对象3D先验。具体来说,用对象识别模型定位任务相关对象,生成相应掩码,用SAM3D提取密集对象级3D表示并与π0的中间视觉特征对齐。模拟实验显示有持续改进,在LIBERO上达到99.1%,在CALVIN上平均长度为4.11。现实世界实验表明该方法在机器人必须跨多个子任务关注不同目标对象的长时操纵场景中特别有效。
英文摘要
Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon $π_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we localize task-relevant objects with object recognition models, generate corresponding object masks, and use SAM3D to extract dense object-level 3D representations, which are aligned with intermediate visual features of $π_0$. This enables the policy to internalize target-object 3D information while preserving the original RGB-language-to-action inference pipeline without requiring depth, point clouds, masks, SAM3D, or additional 3D modules at test time. Simulation experiments show consistent improvements, achieving 99.1\% on LIBERO and an average length of 4.11 on CALVIN. Real-world experiments further demonstrate that our method is particularly effective in long-horizon manipulation scenarios where the robot must focus on different target objects across multiple subtasks.
Comments8 pages, 4 figures