发表机构
University of Waterloo; Soongsil University(滑铁卢大学; 崇实大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出GS-CPE框架,通过由粗到细策略统一几何粗位姿估计与3D高斯溅射精调,在多数据集上实现最优视觉定位性能,兼顾精度与泛化能力。
AI 中文摘要
尽管视觉定位已取得显著进展,从场景坐标回归到直接相机位姿回归,但实现稳健的泛化能力与高精度仍具挑战性。本研究提出GS-CPE(基于高斯溅射的相机位姿估计,Gaussian Splatting based Camera Pose Estimation),一种用于6自由度(6-DoF)相机位姿估计的由粗到细框架,该框架将基于几何的粗位姿估计与基于稳健3D高斯溅射(3DGS)的位姿精调相统一。GS-CPE首先通过在3DGS场景表示上进行检索引导的几何位姿估计来估算粗位姿,随后在多尺度优化框架中通过最小化可见性感知的掩码RGB溅射目标并结合自适应重渲染来精调位姿。在室内和室外基准(包括7Scenes、Cambridge Landmarks、FAST-LIVO2数据集及自定义数据集)上开展的大量实验表明,该方法达到了最优性能,在精度和泛化能力上均持续优于现有方法。
英文摘要
Despite substantial progress in visual localization, from scene coordinate regression to direct camera pose regression, achieving both robust generalization and high accuracy remain challenging. This study introduces GS-CPE (Gaussian Splatting based Camera Pose Estimation), a coarse-to-fine framework for 6-DoF camera pose estimation that unifies geometry-based coarse pose estimation with robust 3D Gaussian Splatting (3DGS) warping based pose refinement. GS-CPE first estimates a coarse pose via retrieval-guided geometric pose estimation on a 3DGS scene representation, then refines it by minimizing a visibility aware masked RGB warping objective in a multi-scale optimization framework, with adaptive re-rendering. Extensive experiments on indoor and outdoor benchmarks including 7Scenes, Cambridge Landmarks, FAST-LIVO2 datasets, and a custom dataset demonstrate state-of-the-art performance, consistently outperforming in both accuracy and generalization.
Comments8 pages, accepted at IROS 2026