DUET-DINO:机器人操作中用于潜在规划的同步跨视角世界建模
DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation
- University of Technology Nuremberg (UTN)(纽伦堡工业大学)
- Technical University of Munich (TUM)(慕尼黑工业大学)
- Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院)
- NVIDIA(英伟达)
- Robotics Institute Germany(德国机器人研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出DUET-DINO,一种同步跨视角潜在世界模型,利用侧视和腕部相机互补信息,在7自由度动作空间上实现潜在规划,在到达、倾斜到达和提升任务上分别取得92%、72.5%和60.0%的成功率,并展现出良好的泛化能力。
AI中文摘要:
动作条件潜在世界模型预测未来的视觉表征,从而实现零样本目标条件机器人规划与控制。然而,它们对细粒度空间和旋转动作的预测对于完整的7自由度末端执行器控制而言并不可靠。为解决这一差距,我们提出了DUET-DINO,一种同步跨视角潜在世界模型,通过跨视角条件化,从静态侧视相机和腕部相机观测中联合学习动作条件预测。通过利用互补的全局场景和夹爪中心信息,DUET-DINO能够在完整的7自由度动作空间上进行潜在规划。在空间多样的到达、方向密集的倾斜到达以及多目标抓取和提升任务中,DUET-DINO始终优于单视角和独立双视角基线,在到达任务中达到92%的成功率,在倾斜到达任务中达到72.5%,在提升任务中达到60.0%。DUET-DINO在DROID和RoboArena数据集上从头训练,并在视觉分布偏移下表现出稳健的泛化能力。我们进一步表明,虽然V-JEPA 2腕部视角预测低估了由细粒度动作引起的视觉动态,但DINOv3预测更好地捕捉了动作条件场景变化,从而带来更强的下游规划能力。代码和模型检查点将开源。项目页面:此https URL。
英文摘要:
Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. To address this gap, we introduce DUET-DINO, a simultaneous cross-view latent world model that jointly learns action-conditioned predictions from static side- and wrist-camera observations through cross-view conditioning. By exploiting complementary global scene and gripper-centric information, DUET-DINO enables latent planning over the full 7-DoF action space. Across spatially diverse reach, orientation-intensive angled-reach, and multi-goal grasp-and-lift tasks, DUET-DINO consistently outperforms single-view and independent dual-view baselines, achieving 92% success on reach, 72.5% on angled-reach, and 60.0% on lift tasks. DUET-DINO is trained from scratch on DROID and RoboArena datasets and generalizes robustly under visual distribution shifts. We further show that while V-JEPA 2 wrist-view predictions underestimate visual dynamics induced by fine-grained actions, DINOv3 predictions better capture action-conditioned scene changes, leading to stronger downstream planning. The code and model checkpoints will be open-sourced. Project page: https://utn-air.github.io/DUET-DINO