NaViRrator:通过习得的视觉路线从人类可读地图进行机器人导航
NaViRrator: Robot Navigation from Human-Readable Maps through a Learned Visual Route
- DGIST(大邱庆北科学技术院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
NaViRRator框架通过SGL-CFM生成路线脚手架,再由视觉语言模型转化为导航指令,实现从人类可读地图到机器人导航,实验证明其成功率与SPL优于现有方法。
AI中文摘要:
人类可读地图为指定机器人目的地提供了直观的界面,但将其示意几何形状与自我中心观察连接起来仍然具有挑战性。我们提出了NaViRRator,一个将用户在此类地图上指定的起点和终点位置转换为预训练的视觉与语言导航(VLN)策略的导航指令的框架。其核心方法RouteScribe将路线推断与语言表达分离,首先在地图图像坐标中生成显式的路线脚手架,然后由预训练的视觉语言模型(VLM)将其转换为导航指令。我们使用起点-终点线条件流匹配(SGL-CFM)构建脚手架,该方法将笔直的起点-终点路点序列变形为地图条件路线。在执行过程中,VLN策略仅接收指令和自我中心观察,而地图和脚手架保持在上游,从而允许在不重新训练地图到语言模块的情况下更换执行器。真实世界实验表明,与直接地图到指令生成、基于A*的脚手架和高斯源条件流匹配相比,成功率和按路径长度加权的成功率(SPL)更高。定性结果进一步显示了更清晰的显著转弯和更好地保留预期机动序列,支持将路线接地语言作为人类可读地图与预训练导航策略之间的模块化接口。
英文摘要:
Human-readable maps provide an intuitive interface for specifying robot destinations, but connecting their schematic geometry to egocentric observations remains challenging. We present NaViRRator, a framework that translates user-specified start and goal locations on such maps into navigation instructions for a pretrained vision-and-language navigation (VLN) policy. Its core method, RouteScribe, separates route inference from verbalization by first generating an explicit route scaffold in map-image coordinates, which a pretrained vision-language model (VLM) converts into a navigation instruction. We construct the scaffold with start--goal line conditional flow matching (SGL-CFM), which deforms a straight start--goal waypoint sequence into a map-conditioned route. During execution, the VLN policy receives only the instruction and egocentric observations, while the map and scaffold remain upstream, allowing executor replacement without retraining the map-to-language modules. Real-world experiments show higher success rates and success weighted by path length (SPL) than direct map-to-instruction generation, A*-based scaffolding, and Gaussian-source conditional flow matching. Qualitative results further show clearer salient turns and better preservation of the intended maneuver sequence, supporting route-grounded language as a modular interface between human-readable maps and pretrained navigation policies.