发表机构
Oxford Brookes University(牛津布鲁克斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对传统视觉SLAM在感知混叠和变化下的脆弱性,本文提出类人语义地点识别方法HuMem-VPR并集成到ORB-SLAM3形成HuMemSLAM,利用自下而上感知与自上而下推理结合,实现高准确率、低延迟的地点识别,显著提升召回率并减少几何后端负担。
AI 中文摘要
自主系统需要可靠的地点识别,以实现高效且有效的同步定位与建图(SLAM)。传统的几何视觉SLAM方法依赖于低级特征和几何一致性,但仍易受感知混叠(不同地点看起来相似)和感知变化(同一地点看起来不同)的影响。尽管语义SLAM和现代学习型视觉地点识别(VPR)方法在具有挑战性的感知条件下提高了鲁棒性,但实时部署需要同时具备高检索准确率和低延迟。受人类记忆和感知的启发,我们提出了HuMem-VPR,它利用自下而上的感知证据与自上而下的上下文推理之间的双向关系来实现高级地点理解。我们进一步引入了HuMemSLAM,即HuMem-VPR与ORB-SLAM3的集成。HuMem VPR在真实图像基准上取得了最高的综合检索准确率,在CARLA基准上取得了具有竞争力的准确率,并且其延迟比所评估的最先进的VPR方法低约两到三倍。在评估的数据集系列和在线实验中,HuMemSLAM相对于ORB-SLAM3的原生检索显著提高了综合Recall @1,同时减少了提交给其几何后端的提议数量。
英文摘要
Autonomous systems require reliable place recognition for efficient and effective simultaneous localisation and mapping (SLAM). Traditional geometric visual SLAM approaches rely on low-level features and geometric consistency, but remain vulnerable to perceptual aliasing, where different places appear similar, and perceptual variation, where the same place appears different. Although semantic SLAM and modern learned visual place recognition (VPR) methods improve robustness under challenging perceptual conditions, real-time deployment requires both high retrieval accuracy and low latency. Inspired by human memory and perception, we propose HuMem-VPR, which exploits the bidirectional relationship between bottom-up perceptual evidence and top-down contextual reasoning to achieve high-level place understanding. We further introduce HuMemSLAM, the integration of HuMem-VPR with ORB-SLAM3. HuMem VPR achieved the highest aggregate retrieval accuracy on the real-image benchmark, competitive accuracy on the CARLA benchmark, and approximately two to three times lower latency than the evaluated state-of-the-art VPR methods. Across the evaluated dataset families and online experiments, HuMemSLAM substantially improved integrated Recall @1 over ORB-SLAM3's native retrieval while reducing the proposals submitted to its geometric backend.
Comments8 pages, 8 figures