发表机构
KTH Royal Institute of Technology; The University of Queensland; University of Technology Nuremberg(瑞典皇家理工学院; 昆士兰大学; 纽伦堡工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出概率场景图(PSG),以高斯层次图表示不确定性并贯穿建图流程,实现实时、高精度的场景图构建与对齐。
AI 中文摘要
3D场景图为机器人感知提供了语义丰富且层次化的表示。然而,现有系统并未将不确定性作为显式信念来维护,也未通过构建和细化图的操作来传播不确定性。我们引入了概率场景图(PSG),这是传统场景图的一种泛化,它表示可能图上的后验分布,分解为实体、关系和语义属性的离散图结构,以及对其进行空间锚定的连续状态,并对这两个组成部分都维护不确定性。几何信息直接由节点承载,而不是从单独构建的度量地图中选择,因此,在需要时,度量地图由图生成,而非先于图存在。我们使用高斯层次图(HGG)实例化PSG的概率空间锚定:每个对象基元在正态-逆威沙特信念下由全协方差高斯表示,同一参数化在节点内递归应用,产生一个以更精细分辨率解析其表面的几何图。然后,我们构建了一个映射流程,在整个图构建和细化过程中保持这些信念:一种纯基于图的从粗到细的对齐通过比较节点信念来配准观测,而嵌套的期望最大化与因子图优化联合细化位姿、对象参数和内部几何。在涵盖室内RGB-D、室外LiDAR和跨模态部署的六个数据集上,HGG以传感器速率运行,内存近乎恒定,并实现了最先进的对象精度和零样本图对齐。
英文摘要
3D scene graphs provide semantically rich and hierarchical representations for robot perception. However, existing systems do not maintain uncertainty as an explicit belief or propagate it through the operations that construct and refine the graph. We introduce Probabilistic Scene Graph (PSG), a generalization of the conventional scene graph that represents a posterior over possible graphs, factorized into a discrete graph structure of entities, relations, and semantic attributes, and continuous states that ground them spatially, with uncertainty maintained over both components. Geometry is carried directly by the nodes rather than selected from a separately constructed metric map, so a metric map, where needed, follows from the graph rather than preceding it. We instantiate PSG's probabilistic spatial grounding with hierarchical graphs of Gaussians (HGG): each object primitive is represented by a full-covariance Gaussian under a Normal-Inverse-Wishart belief, and the same parametrization applied recursively within a node yields a geometry graph that resolves its surface at finer resolution. We then build a mapping pipeline that preserves these beliefs throughout graph construction and refinement: a purely graph-based coarse-to-fine alignment registers observations by comparing node beliefs, while a nested Expectation-Maximization and factor-graph optimization jointly refines poses, object parameters, and internal geometry. Across six datasets spanning indoor RGB-D, outdoor LiDAR, and cross-modality deployment, HGG operates at sensor rate with near-constant memory and achieves state-of-the-art object accuracy and zero-shot graph alignment.