arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoSe-SLAM:面向动态单目视频的鲁棒语义感知高斯溅射SLAM

RoSe-SLAM: Robust Semantic-Aware Gaussian Splatting SLAM from Dynamic Monocular Videos

Wenting Wang, Jiaxin Guo, Wenzhen Dong, Yun-Hui Liu, Charlie C. L. Wang, Yeung Yam

arXiv 2608.29003首次发表:更新:

发表机构

The Chinese University of Hong Kong; The University of Manchester; Centre for Perceptual and Interactive Intelligence (CPII) Limited(香港中文大学; 曼彻斯特大学; 感知与互动智能中心有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出RoSe-SLAM,利用2D基础模型语义特征与时空运动掩码等模块,结合几何与语义线索,在动态单目视频中实现更优的SLAM性能,优于现有动态RGB SLAM基线。

AI 中文摘要

在动态且非结构化的环境中,传统SLAM系统因采用静态假设,通常会出现显著的精度退化问题。本研究提出鲁棒语义感知高斯溅射SLAM(RoSe-SLAM),旨在通过对未标定单目输入进行整体语义场景理解,解决动态挑战,实现精确的相机跟踪与高质量几何重建。与采用手工语义标签的传统语义SLAM不同,RoSe-SLAM利用二维基础模型的语义特征提升动态跟踪与建图性能,通过将丰富的语义特征蒸馏到高斯场中,有效识别动态干扰项并实现语义感知的多视图一致性,显著增强几何重建与场景修复能力。具体而言,本研究提出时空运动掩码生成模块,可同时实现长期运动监测与短期瞬态动态捕捉,实现动态物体与静态背景的鲁棒且有效的解耦。在全局光束平差过程中,提出基于遮挡的关键帧选择机制,以遮挡程度作为关键帧选择的度量标准,同时提出多视图语义一致性模块以提升动态环境下的建图质量。通过将几何运动线索与语义先验相结合,系统可动态过滤不可靠观测结果并重建精确的静态场景几何。在动态TUM、Bonn和Wild-Mocap数据集等基准数据集上开展的大量实验表明,所提方法在轨迹估计与静态场景建图两方面均实现了优异性能,在长期动态室内环境中优于现有动态RGB SLAM基线方法。

英文摘要

In dynamic and unstructured environments, conventional SLAM systems generally suffer from significant accuracy degeneration due to their static assumptions. In this work, we propose Robust Semantic-aware Gaussian Splatting SLAM (RoSe-SLAM), to address the dynamic challenge by a holistic semantic scene understanding from uncalibrated monocular inputs, achieving accurate camera tracking and high-quality geometry reconstruction. Unlike conventional semantic SLAM using handcrafted semantic labels, our RoSe-SLAM exploits the semantic feature from 2D foundation model to enhance the dynamic tracking and mapping performance. By distilling the rich semantic features to our Gaussian fields, our method effectively identifies dynamic distractors and achieves semantic-aware multi-view consistency, significantly enhancing the geometric reconstruction and scene inpainting. Specifically, we propose a spatial-temporal motion mask generation module, enabling both long-term motion monitoring and short-term transient dynamics capturing, achieving robust and effective disentanglement of dynamic objects and static backgrounds. During global bundle adjustment, we propose an occlusion-aware keyframe selection mechanism to prioritize the occlusion as metric to pick the keyframes, and a multi-view semantic consistency module to improve the mapping quality in dynamic environments. By combining geometric motion cues with semantic priors, our system dynamically filters unreliable observations and reconstructs accurate static scene geometry. Extensive experiments conducted on benchmark datasets including dynamic TUM, Bonn and Wild-Mocap datasets, demonstrate that our method achieves superior performance in both trajectory estimation and static scene mapping, outperforming existing dynamic RGB SLAM baselines in long-term dynamic indoor environments.

CommentsAccepted by IEEE/RSJ INTERNATIONAL CONFERENCE ON INTELLIGENT ROBOTS & SYSTEMS (IROS), 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑