MVG-WAM:用于机器人操作的多视图几何感知世界-动作建模
MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation
浏览论文内容
中文总结 AI 辅助
提出MVG-WAM模型,通过极线约束全局状态与视图索引几何状态结合,并采用多时域深度监督,在LIBERO、RoboTwin 2.0及真实机器人上分别取得99.1%、92.07%和91.3%的成功率。
中文摘要 AI 辅助
世界-动作模型(WAMs)将视觉动力学与动作预测相结合,将预训练视频模型的丰富先验知识引入机器人操作中。然而,其多视图接口通常采用图像拼接或令牌连接的方式,使得同步相机之间的几何关系隐含不清。这使得将全局场景上下文与交互所需的局部几何信息联系起来变得更加困难。我们提出了多视图几何感知世界-动作模型(MVG-WAM),该模型将这些观测组织为同一物理世界的相关投影,而非画布上的独立图像。我们的模型结合了极线约束的全局状态与从同步观测中联合推断的视图索引几何状态。相机感知路由为每个视频区域提供相应的几何上下文和共享的全局状态,显式地构建用于动作预测的表示。我们进一步通过多时域未来深度监督将几何感知表示锚定在度量尺度上,而无需在动作展开过程中进行深度解码。MVG-WAM在LIBERO基准上平均成功率达99.1%,在RoboTwin 2.0上达92.07%,在两个基准上均展现出具有竞争力的性能。在Cobot Magic上的真实世界实验进一步表明,在涵盖三项操作任务的150次试验中,成功率达91.3%。
英文摘要
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it harder to connect global scene context with the local geometry required for interaction. We introduce the Multi-View Geometry-Aware World-Action Model (MVG-WAM), which organizes these observations as related projections of one physical world rather than separate images on a canvas. Our model combines an epipolar-constrained global state with view-indexed geometric states jointly inferred from synchronized observations. Camera-aware routing supplies each video region with its corresponding geometric context and the shared global state, explicitly structuring the representation used for action prediction. We further ground the geometry-aware representation in metric scale through multi-horizon future-depth supervision, without requiring depth decoding during action rollout. MVG-WAM achieves average success rates of 99.1% on LIBERO and 92.07% on RoboTwin 2.0, demonstrating competitive performance across both benchmarks. Real-world experiments on Cobot Magic further demonstrate a 91.3% success rate across 150 trials spanning three manipulation tasks.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- The Hong Kong University of Science and Technology(香港科技大学)
- ETH Zurich(苏黎世联邦理工学院)
- Zhejiang University(浙江大学)
- EPFL(洛桑联邦理工学院)
- University of Zurich(苏黎世大学)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。