发表机构
KAIST AI; Holiday Robotics(韩国科学技术院人工智能研究所; 假日机器人公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉-语言-动作模型中因观察与动作定义的帧不匹配问题,提出以机器人为中心的点图方法,该方法能保留2D VLA所需网格,以最小架构变化集成到现有模型,实验证明其在模拟和实际机器人实验中表现优于基线。
AI 中文摘要
视觉-语言-动作(VLA)模型根据视觉观察和语言指令预测机器人动作。动作在机器人自身的3D坐标系中定义,但大多数VLA在相机帧中观察场景,导致观察场景的位置与定义动作的位置之间存在帧不匹配。在固定视点下这种不匹配影响较小,而随着大规模数据集聚合不同相机设置下的演示,且策略必须跨视点进行泛化时,问题变得更严重。我们以机器人为中心的点图来解决这种不匹配,其像素存储机器人帧中场景点的3D坐标。点图提供机器人帧3D几何结构,同时保留预训练2D VLA所需的密集H x W网格,能以最小架构变化集成到现有VLA中。在RoboCasa上,点图改进了pi0.5和SmolVLA,优于代表性的相机视点和3D感知基线。在实际机器人实验中,当相机移到训练中未见过的位置时,相对于仅使用RGB的策略,点图优势更明显。
英文摘要
Vision-language-action (VLA) models require 3D spatial reasoning, yet RGB observations encode robot-object geometry only implicitly. Lifting depth with camera intrinsics makes this geometry explicit as dense, image-aligned pointmaps, but their camera-frame coordinates depend on camera placement. We propose SeeR-VLA, which transforms pointmaps into a robot-centric frame with an end-effector origin and robot-base-aligned axes. An encoder initialized from pretrained RGB weights extracts pointmap features, which are added to corresponding RGB tokens without increasing the token count. Across 24 RoboCasa tasks and four real-world tasks, SeeR-VLA improves average success over RGB-only $π_{0.5}$ by 6.4 and 32.5 percentage points, respectively. It also exceeds the strongest evaluated 3D-augmented baseline, PointVLA, by 3.5 and 23.7 percentage points, respectively. Beyond these gains, our ablations clarify how coordinate choices affect VLA performance, showing that end-effector centering is most effective with robot-base-aligned axes. The benefits grow as training viewpoints diversify, highlighting the importance of using robot-frame pointmaps when learning from diverse camera configurations.
CommentsProject page: https://davian-robotics.github.io/pointmap/