arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19824cs.RO

TADreamer:面向陆空双模态机器人的零样本语言引导三维导航——基于视频想象

TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination

Xiangyu Li, Tiancheng Lai, Xijie Huang, Ruitian Pang, Siqi Shen, Juncheng Chen, Zaisheng Pan, Chao Xu, Fei Gao, Yanjun Cao

首次发表
浏览论文内容

中文总结 AI 辅助

TADreamer提出零样本框架,利用视频想象与两阶段标定实现陆空双模态机器人语言引导导航,在真实场景中显著降低深度误差。

中文摘要 AI 辅助

面向陆空双模态机器人的语言引导导航需要选择与场景上下文和任务意图相匹配的路线及运动模态。生成的视频能够表征此类运动序列,但由于尺度模糊性和轴相关的几何畸变,从中恢复度量一致的导航参考具有挑战性。我们提出TADreamer,一种零样本框架,无需任务特定训练或微调即可将视频想象的导航锚定于实测几何。视觉-语言模型将机载观测和指令转化为导航提示,筛选有效的生成视频,并在需要重新生成时提供纠正性反馈。所选视频被重建为标注陆空模态的三维航点。两阶段标定流程利用视场约束初始化尺度估计,随后通过将重建点云配准至实测几何来细化轴相关尺度、旋转和平移。标定后的航点及模态标签引导规划器,该规划器融合实测几何以执行机器人动作。真实世界实验展示了在七个室内外场景中的导航能力。在每轮五个候选视频的情况下,所有七个场景均在两轮内获得可用视频。在标定观测上,与NavDreamer相比,我们的方法将平均绝对深度误差降低了87.7%,平均绝对相对深度误差降低了86.3%。

英文摘要

Language-guided navigation for terrestrial-aerial bimodal robots requires selecting routes and locomotion modes that match scene context and task intent. Generated videos can represent such motion sequences, but recovering metrically consistent navigation references from them is challenging because of scale ambiguity and axis-dependent geometric distortions. We present TADreamer, a zero-shot framework that grounds video-imagined navigation in measured geometry without task-specific training or fine-tuning. A vision-language model translates onboard observations and instructions into navigation prompts, selects valid generated videos, and provides corrective feedback when regeneration is needed. The selected video is reconstructed into 3D waypoints annotated with terrestrial or aerial modes. A two-stage calibration procedure uses field-of-view constraints to initialize scale estimation, then refines axis-dependent scales, rotation, and translation by registering the reconstructed point cloud to measured geometry. The calibrated waypoints and mode labels guide a planner that incorporates measured geometry for robot execution. Real-world experiments demonstrate navigation across seven indoor and outdoor scenarios. With five candidates per round, usable videos are obtained within two rounds in all seven scenarios. On the calibration observations, our method reduces mean absolute depth error by 87.7% and mean absolute relative depth error by 86.3% compared with NavDreamer.

发表机构

  • Zhejiang University(浙江大学)
  • Huzhou Institute of Zhejiang University(浙江大学湖州研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑