发表机构
Uni-Ubi(优必选(Uni-Ubi))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出LaTraNav框架,结合慢速VLM与快速流匹配规划器,通过仿真数据生成流水线训练,实现语言引导的自适应视觉导航,提升了规划性能与路径更新率。
AI 中文摘要
可通行性是视觉导航的核心要素,但会随机器人能力和用户偏好变化。传统流程常依赖显式代价地图或带预定义准则的分割掩码,需手工规则与精细调参;且视角依赖的分割掩码会加剧感知延迟下的异步规划难题。本文提出LaTraNav框架,用于学习语言条件化的可通行性表征以实现自适应视觉导航。其异步架构结合了慢速视觉语言模型(VLM),该模型生成可通行性与导航目标的潜在表征,以及基于这些表征的快速流匹配规划器。为训练该系统,我们开发了带可控轨迹的仿真数据生成流水线,生成与语言指令、可通行性地图、目标位置及多样轨迹配对的观测; photorealistic图像翻译进一步提升视觉真实性。对多源数据集的评估显示,慢速VLM可实现有效的语言引导可通行性分割与目标定位,快速规划器则完成自适应像素空间路径规划;潜在条件化相比显式分割掩码提升了规划性能,而异步调度在相同语义更新率下使路径更新率提升6.05倍。
英文摘要
Traversability is essential for visual navigation but varies with robot capabilities and user preferences. Conventional pipelines often rely on explicit costmaps or segmentation masks with predefined criteria, requiring hand-crafted rules and careful tuning. Moreover, viewpoint-dependent segmentation masks complicate asynchronous planning under perception latency. We present LaTraNav, a framework that learns language-conditioned traversability representations for adaptive visual navigation. Its asynchronous architecture combines a slow vision-language model that produces latent representations of traversability and navigation goals, with a fast flow-matching planner conditioned on these representations. To train the system, we develop a simulation-based data generation pipeline with controllable trajectories, producing observations paired with language instructions, traversability maps, goal locations, and diverse trajectories. Photorealistic image translation further enhances visual realism. Evaluations on datasets from multiple sources demonstrate effective language-guided traversability segmentation and goal localization by the slow VLM, alongside adaptive pixel-space path planning by the fast planner. Latent conditioning improves planning performance over explicit segmentation masks, while asynchronous scheduling increases the path-update rate by $6.05\times$ at the same semantic-update rate.