MASt3R-Fusion:融合前馈视觉模型与IMU、GNSS的高功能SLAM系统
MASt3R-Fusion: Integrating Feed-Forward Visual Model with IMU, GNSS for High-Functionality SLAM
- School of Geodesy and Geomatics, Wuhan University(武汉大学测绘学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有前馈点图回归SLAM未融合多传感器信息的不足,提出紧耦合IMU、GNSS的MASt3R-Fusion框架,通过分层因子图实现实时高精度全局一致建图,性能优于现有同类系统。
AI中文摘要:
视觉SLAM是机器人、自动驾驶与扩展现实(XR)领域的基石技术,但传统系统往往在低纹理环境、尺度模糊问题以及挑战性视觉条件下出现性能退化。近期基于前馈神经网络的点图回归技术取得进展,可直接从图像恢复高保真3D场景几何结构,依托学习到的空间先验克服传统多视图几何方法的局限,但这类流程通常舍弃了经广泛验证的概率多传感器信息融合优势。本研究提出MASt3R-Fusion,这是一款多传感器辅助的视觉SLAM框架,将前馈点图回归与惯性测量单元(IMU)数据、全球导航卫星系统(GNSS)数据等互补传感器信息紧耦合。该系统将基于Sim(3)的视觉对齐约束(以Hessian形式)引入通用度量尺度SE(3)因子图以实现高效信息融合,设计了分层因子图结构,可同时支持实时滑动窗口优化与带激进回环检测的全局优化,实现实时位姿跟踪、度量尺度结构感知与全局一致建图。我们在公开基准数据集与自采数据集上评估了所提方法,结果表明其相比现有以视觉为中心的多传感器SLAM系统在精度与鲁棒性上均有大幅提升。代码将开源以支持结果复现与后续研究(https://github.com/GREAT-WHU/MASt3R-Fusion)。
英文摘要:
Visual SLAM is a cornerstone technique in robotics, autonomous driving and extended reality (XR), yet classical systems often struggle with low-texture environments, scale ambiguity, and degraded performance under challenging visual conditions. Recent advancements in feed-forward neural network-based pointmap regression have demonstrated the potential to recover high-fidelity 3D scene geometry directly from images, leveraging learned spatial priors to overcome limitations of traditional multi-view geometry methods. However, the widely validated advantages of probabilistic multi-sensor information fusion are often discarded in these pipelines. In this work, we propose MASt3R-Fusion,a multi-sensor-assisted visual SLAM framework that tightly integrates feed-forward pointmap regression with complementary sensor information, including inertial measurements and GNSS data. The system introduces Sim(3)-based visualalignment constraints (in the Hessian form) into a universal metric-scale SE(3) factor graph for effective information fusion. A hierarchical factor graph design is developed, which allows both real-time sliding-window optimization and global optimization with aggressive loop closures, enabling real-time pose tracking, metric-scale structure perception and globally consistent mapping. We evaluate our approach on both public benchmarks and self-collected datasets, demonstrating substantial improvements in accuracy and robustness over existing visual-centered multi-sensor SLAM systems. The code will be released open-source to support reproducibility and further research (https://github.com/GREAT-WHU/MASt3R-Fusion).