无序地标视觉导航
Unordered Landmark Visual Navigation
浏览论文内容
中文总结 AI 辅助
针对无序图像导航的挑战,本文提出无需时间与里程计先验的ULVN框架,通过整合建图、定位与规划,在仿真和真实场景中性能优于现有最优方法。
中文摘要 AI 辅助
图像目标导航是具身AI的基础能力,但其实际部署受限于强先验假设。现有方法主要依赖时间有序视频流或辅助传感器(如深度、LiDAR)维持空间一致性,这些序列与多模态依赖严重限制了可扩展性,尤其在使用众包或预记录的无序图像集合部署机器人时。当时间先验被移除,现有方法会遭遇严重感知混淆、噪声关联和灾难性建图失败。为解决这一未被充分探索的挑战,我们提出无序地标视觉导航(Unordered Landmark Visual Navigation,ULVN),一个仅使用RGB的统一框架,无需时间和里程计先验。ULVN通过整合建图、定位与规划系统性缓解误差累积:具体而言,它通过校准几何验证和最大生成树优化,直接从非结构化图像构建鲁棒的2D拓扑图;在闭环执行中,ULVN摒弃序列启发式方法,采用基于图的置信传播滤波器结合熵自适应融合,实现全局定位与动态子目标规划。大量仿真与真实部署实验表明,ULVN的性能显著优于现有最优方法。
英文摘要
Image-goal navigation is a fundamental capability for embodied AI, yet its practical deployment is strained by strong prior assumptions. Existing methods predominantly rely on temporally ordered video streams or auxiliary sensors (e.g., depth, LiDAR) to maintain spatial consistency. These sequential and multimodal dependencies severely restrict scalability, especially when deploying robots using crowd-sourced or pre-recorded unordered image collections. When temporal priors are removed, current methods struggle with severe perceptual aliasing, noisy associations, and catastrophic mapping failures. To address this underexplored challenge, we propose Unordered Landmark Visual Navigation (ULVN), a unified RGB-only framework free from temporal and odometric priors. ULVN systematically mitigates error accumulation by integrating mapping, localization, and planning. Specifically, it constructs a robust 2D topological map directly from unstructured images via calibrated geometric verification and maximum spanning forest refinement. For closed-loop execution, ULVN abandons sequential heuristics, utilizing a graph-based belief propagation filter with entropy-adaptive fusion for global localization and dynamic subgoal planning. Extensive experiments in simulation and real-world deployments demonstrate that ULVN significantly outperforms state-of-the-art methods.
发表机构
- Sun Yat-sen University(中山大学)
- Insta360 Research(影石创新研究院)
机构由 AI 辅助整理,请以论文原文为准。