arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Prior-SG:面向任意结构环境中场景图的任务与先验驱动区域分割

Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments

Giorgio Tonetti, Laurent Kneip, Abel Gawel, Marco Hutter

arXiv 2608.06170首次发表:更新:

发表机构

RAI Institute; ETH Zurich(RAI研究所; 苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Prior-SG框架,将场景图生成视为概率对齐问题,通过融合异构专家与拓扑先验提升任意结构环境的语义区域分割精度,实现零样本本体灵活性。

AI 中文摘要

分层三维场景图是自主移动平台高层空间推理的一种有前景的表示方式。然而,现有的提取框架通常依赖纯局部视觉聚类或严格的几何启发式方法(如墙体分隔的房间),在开放布局或任意结构环境中表现不佳。我们提出Prior-SG,这是一种任务与先验驱动的框架,从根本上将场景图生成视为概率对齐问题。当机器人探索时,它会利用多尺度开放词汇特征融合策略,不断将传入的RGB-D传感器流聚合为物理接地的实例图。随后,系统通过最大后验概率(MAP)估计推断该地图的高层功能语义,该估计由先验图引导——先验图是大型语言模型动态合成的环境结构与任务相关词汇的逻辑预期。通过优化马尔可夫随机场(MRF),将异构专家(视觉、几何和离散对象)与这些拓扑先验融合,系统可解决局部感知歧义。我们在多样化的模拟住宅数据集和大型开放布局真实环境中验证了该方法。与近期基线相比,Prior-SG实现了最先进的语义区域分割精度,在无实体墙体时能稳健地描绘远处的功能边界,且具有独特的零样本本体灵活性,使机器人能根据给定的高层任务完全重构其空间划分。

英文摘要

Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or strict geometric heuristics, such as wall-separated rooms, which fail in open-plan or arbitrarily-structured environments. We propose Prior-SG, a task- and prior-driven framework that casts scene graph generation fundamentally as a probabilistic alignment problem. As the robot explores, it continuously aggregates an incoming RGB-D sensor stream into a physically grounded Instance Graph utilizing a multi-scale, open-vocabulary feature fusion strategy. The system then infers the high-level functional semantics of this map through a Maximum A Posteriori (MAP) estimate, guided by a Prior Graph-a logical expectation of the environment's structure and task-relevant vocabulary synthesized dynamically by a Large Language Model. By optimizing a Markov Random Field that fuses heterogeneous experts (visual, geometric, and discrete objects) with these topological priors, the system resolves local perceptual ambiguities. We validate this approach across diverse simulated residential datasets and large, open-plan real-world environments. Prior-SG achieves state-of-the-art semantic region segmentation accuracy compared to recent baselines, robustly delineates distant functional boundaries in the absence of physical walls, and uniquely provides zero-shot ontological flexibility, enabling the robot to entirely restructure its spatial partitioning based on a given high-level task.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑