arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

M2-SMap:基于分层多模型表示的内存高效语义建图

M2-SMap: Memory-Efficient Semantic Mapping with Hierarchical Multi-Model Representation

QiYing Deng, ZhongLai Wang, Yuan Gao, Wei Dong

arXiv 2608.07074首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Shanghai Jiao Tong University(电子科技大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于分层多模型表示的内存高效语义建图框架M2-SMap,通过分层几何分解等策略提升场景表达,在三个RGB-D序列上实现低基元数、无物体粘连的实时建图。

AI 中文摘要

密集点云图作为常用的建图表示,因内存消耗随场景规模快速增长,难以部署在资源受限的机器人上。尽管紧凑的单模型表示可降低内存成本,但其固定的几何表达能力不足以应对结构多样的环境。现有多模型方法提升了表达灵活性,但其特征提取与模型选择常受局部几何主导,易导致过拟合及物体间粘连。为解决这些问题,本文提出M2-SMap,一种基于分层多模型表示的内存高效语义建图框架。首先,分层几何分解将RGB-D点云划分为紧凑的高斯分量;随后,投影引导的语义标注机制为每个分量分配实例标识;接着,将这些标注融入面向物体的高斯融合策略;此外,多尺度特征提取策略分离出大型平面区域、语义物体及复杂残差结构,分别用有界平面、物体级超二次曲面及高斯混合模型(GMM)基元表示。在三个RGB-D序列上的实验表明,M2-SMap以不低于29.37Hz的帧率实时运行,且基元数量最少,较最优基准方法平均减少18.7%;还将每帧测量的物体间粘连平均数量从2.808降至0,实现了高效且语义一致的场景表示。

英文摘要

Dense point cloud maps, as a typically used mapping representation, are difficult to deploy on resource-constrained robots because their memory consumption grows rapidly with scene scale. Although compact single-model representations reduce memory cost, their fixed geometric expressiveness is insufficient for structurally diverse environments. Existing multi-model methods improve representational flexibility, yet their feature extraction and model selection are often dominated by local geometry, which can cause overfitting and adhesion between objects. To address these issues, this paper presents M2-SMap, a memory-efficient semantic mapping framework based on hierarchical multi-model representation. First, a hierarchical geometric decomposition partitions RGB-D point clouds into compact Gaussian components. Then, a projection-guided semantic annotation mechanism assigns instance identities to each component. Subsequently, these annotations are incorporated into an object-aware Gaussian fusion strategy. Furthermore, a multi-scale feature extraction strategy separates large planar regions, semantic objects, and complex residual structures, which are respectively represented by bounded planes, object-level superquadrics, and GMM primitives. Experiments on three RGB-D sequences show that M2-SMap runs in real time at no less than 29.37 Hz while achieving the lowest primitive count, with an average reduction of 18.7% over the best baseline. It also reduces the mean per-frame number of measured inter-object adhesion cases from 2.808 to 0, demonstrating efficient and semantically consistent scene representation.

Comments8 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑