发表机构
University of Zaragoza; German Aerospace Center (DLR); Karlsruhe Institute of Technology (KIT)(萨拉戈萨大学; 德国航空航天中心; 卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对类行星地形中视觉定位难题,提出基于 3D 高斯溅射构建可微地图表示估计相机姿态的方法,采用结合光度与几何损失的几何感知训练策略,经实验验证能有效提升定位准确性与姿态估计鲁棒性。
AI 中文摘要
在类行星地形中,视觉定位极具挑战性,其特点是纹理低、感知别名、光照 harsh、前向漫游车运动及无约束行驶方向导致的稀疏、弱重叠视图。当前最先进的图像到图像和图像到地图匹配管道性能显著下降。本文提出一种视觉重定位方法,直接针对用 3D 高斯溅射构建的可微地图表示估计相机姿态,不同于传统基于对应关系的管道。关键贡献是几何感知训练策略,结合光度和几何损失,首次通过结合多视图立体和激光雷达深度提供几何监督。实验表明联合优化产生更适合场景几何的 3DGS 模型,提高了光度和几何一致性及单图像 6 自由度姿态估计的鲁棒性和准确性。在行星模拟环境数据上的大量实验验证了该方法的有效性。
英文摘要
Visual localization becomes extremely challenging in planetary-like terrains characterized by low texture, perceptual aliasing, harsh illumination, and sparse, weakly overlapping viewpoints induced by forward rover motion and unconstrained driving directions. Under these conditions, state-of-the-art image-to-image and image-to-map matching pipelines suffer significant performance degradation. In this work, we propose a visual relocalization method that departs from classical correspondence-based pipelines by directly estimating camera poses against a differentiable map representation built with 3D Gaussian Splatting (3DGS). Our key contribution is a geometry-aware training strategy that combines photometric and geometric losses, where the geometric supervision is provided for the first time by combining multi-view stereo (MVS) and LiDAR depths. We show that this joint optimization produces a 3DGS model that better fits the underlying scene geometry, leading to improved photometric and geometric consistency and more robust, accurate single-image 6-DoF pose estimation. Extensive experiments on data acquired in planetary-analog environments validate the effectiveness of our approach, showing substantial gains in relocalization accuracy under challenging conditions. Code is available at https://github.com/DLR-RM/multimodal-gsplat-relocalization.
CommentsAccepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)