发表机构
Carnegie Mellon University; NVIDIA(卡内基梅隆大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GeomVLA通过共享3D坐标系统一感知、运动预测和动作生成,利用轨迹去噪器提取运动标记并条件化动作去噪器,在CALVIN上取得最优性能,强调几何一致性的关键作用。
AI 中文摘要
我们提出了GeomVLA,一种视觉-语言-动作(VLA)模型,它在共享的以机器人为中心的3D坐标框架内统一了感知、潜在场景运动预测和动作生成。我们的方法利用深度和相机标定将预训练的VLM特征提升为空间上有依据的3D场景标记,同时保留在VLM预训练期间学到的语义表示。我们进一步引入了3D场景轨迹去噪器,这是一个任务条件模块,学习场景点在3D中预期如何移动的潜在表示。GeomVLA不是将预测的轨迹作为开环计划执行,而是从轨迹去噪器中提取中间运动标记,并通过几何感知注意力将其用于条件化基于3D流的动作去噪器。GeomVLA在CALVIN上达到了最先进的性能,在LIBERO和RoboTwin2.0上表现具有竞争力,并且在没有机器人动作预训练的情况下,在真实世界操作设置中优于强基线。大量消融实验表明,仅靠未来运动推理是不够的:主要收益与在整个感知到动作流程中保持场景表示、运动预测和机器人动作之间的几何一致性有关。
英文摘要
We present GeomVLA, a Vision-Language-Action (VLA) model that unifies perception, latent scene motion prediction, and action generation within a shared robot-centric 3D coordinate frame. Our approach lifts pretrained VLM features into spatially grounded 3D scene tokens using depth and camera calibration, while retaining the semantic representations learned during VLM pretraining. We further introduce a 3D Scene Trajectory Denoiser, a task-conditioned module that learns a latent representation of how scene points are expected to move in 3D. Rather than executing the predicted trajectory as an open-loop plan, GeomVLA extracts intermediate motion tokens from the trajectory denoiser and uses them to condition a 3D flow-based action denoiser through geometry-aware attention. GeomVLA achieves state-of-the-art performance on CALVIN, competitive performance on LIBERO and RoboTwin2.0, and outperforms strong baselines in real-world manipulation settings without robot-action pretraining. Extensive ablations show that future-motion reasoning alone is insufficient: the primary gains are associated with maintaining geometric consistency among scene representation, motion prediction, and robot actions throughout the perception-to-action pipeline.
CommentsAccepted to CoRL 2026. Project page: https://ziyin-xiong.github.io/geomvla.io/