arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2609.16737cs.ROcs.AIcs.CVcs.LG

看见关键:视觉提示引导的视频规划用于可泛化机器人导航

Visual Cue Guided Video Planning for Generalizable Robot Navigation

  • Technical University of Munich(慕尼黑工业大学)
  • Massachusetts Institute of Technology(麻省理工学院)
  • Örebro University(厄勒布鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

Hojin Lee, Sizhe Lester Li, Maximilian Hilger, Susie Lu, Achim J. Lilienthal, Vincent Sitzmann, Daniel A. Duecker

AI总结:

提出CueNav,结合视觉提示引导的视频规划与逆动力学模型,实现长视界、具身感知的泛化机器人导航,显著提升迷宫和狭窄通道成功率。

AI中文摘要:

生成式视频模型通过预测未来观测作为视频计划,可成为机器人导航的有力骨干。近期方法通常将视频规划条件化为短视界引导,并通过场景重建恢复几何路点,对更长视界规划和精确的视频到动作转换探索较少。我们提出CueNav,一种基于视频模型的导航框架,结合视觉提示引导的视频规划与具身特定的逆动力学模型(IDM)。作为视觉提示,我们使用鸟瞰图(BEV)地图传达全局任务上下文,并在自我中心观测中保留部分机器人身体以暴露具身上下文。这些提示引导视频规划器,而IDM将视频计划中提取的稠密光流场转换为机器人动作。凭借编码全局任务上下文的视觉提示,CueNav在迷宫导航中的成功率比无提示规划高出近2倍。具身感知视图与IDM使得在狭窄通道中实现70%的成功率精确导航,而对比方法大多无法完成任务。我们进一步展示了零样本语义条件导航以及同一视频规划器在不同机器人平台上的部署。我们的结果表明,视觉提示引导的视频规划与具身特定动作接地为面向更长视界规划和具身感知控制的泛化导航框架铺平了道路。更多结果和代码可在我们的项目网站获取:此https URL。

英文摘要:

Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guidance and recover geometric waypoints through scene reconstruction, leaving longer-horizon planning and precise video-to-action translation less explored. We present CueNav, a video model-based navigation framework combining visual cue guided video planning with an embodiment-specific Inverse-Dynamics Model (IDM). As visual cues, we use a Bird's-Eye View (BEV) map to convey global task context and retain part of the robot body in the egocentric observation to expose embodiment context. These cues guide the video planner, while the IDM translates dense flow fields extracted from the video plan into robot actions. With the visual cue encoding global task context, CueNav achieves nearly 2x higher success in maze navigation than planning without the cue. The body-aware view with the IDM enables precise navigation with 70% success in a narrow passage where comparison methods largely fail to complete the task. We further demonstrate zero-shot semantic-conditioned navigation and deployment of the same video planner across different robot platforms. Our results show that visual cue-guided video planning with embodiment-specific action grounding paves the way toward a generalizable navigation framework for longer-horizon planning and embodiment-aware control. Additional results and code are available on our project website: https://cuenav.github.io.

补充信息

↑