arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22351cs.CV

三维场景图中语义不确定性的层次聚合

Hierarchical Aggregation of Semantic Uncertainty in 3D Scene Graphs

Carlos Cueto Zumaya, Iacopo Catalano, Wallace Moreira Bessa, Julio A. Placed

首次发表
浏览论文内容

中文总结 AI 辅助

针对开放词汇三维场景图中条目不确定性未区分的问题,提出利用检测器置信度和嵌入,通过层次传播概率,在无额外训练下提高对象检索并降低房间级断言错误。

中文摘要 AI 辅助

开放词汇三维场景图(3DSGs)将每个对象节点锚定在视觉-语言嵌入中,然而它们将每个条目记录为同等确定,因此查询地图的机器人无法分辨哪些条目不可靠。语义不确定性的估计器可以提供这种区分,但它们需要重复采样模型、训练或保留标签,而这些在部署系统查询时均不可用。我们提出一个框架,利用检测器置信度和3DSG已存储的嵌入,将其转换为条目正确的概率,并通过包含层次将该概率传播到房间包含查询类别的信念中。四个信号,每个信号与它所指示的对象级错误配对,在视觉-语言模型学习的对数尺度上转换为概率,并以封闭形式组合,无需额外的感知或训练。共享检测器和词汇表的对象会一起失败,因此框架在完全相关极限下聚合它们,而在独立假设下的聚合会将一次重复错误视为重复证据。在HM3DSem上针对最先进的三维场景图系统进行评估,该框架提高了对象检索性能,并降低了其读取图的房间级断言的错误。

英文摘要

Open-vocabulary 3D Scene Graphs (3DSGs) ground each object node in a vision-language embedding, yet they record every entry as equally certain, so a robot querying the map cannot tell which of its entries are unreliable. Estimators of semantic uncertainty could supply that distinction, but they require repeated sampling of a model, training, or held-out labels, none of which are available to a deployed system at query time. We present a framework that exploits the detector confidence and the embeddings a 3DSG already stores, converts them into a probability that an entry is correct, and propagates that probability through the containment hierarchy into a belief that a room contains a queried class. Four signals, each paired with the object-level error it indicates, are converted to probabilities at the logit scale learned by the vision-language model and combined in closed form with no additional perception or training. Objects sharing a detector and a vocabulary fail together, so the framework aggregates them in the fully correlated limit, where an aggregation under independence would treat one repeated error as repeated evidence. Evaluated on HM3DSem against a state-of-the-art 3DSG system, the framework improves object retrieval and lowers the error of the room-level assertions of the graph it reads.

发表机构

  • University of Turku(图尔库大学)
  • Aragón Institute of Technology (ITA)(阿拉贡技术研究所)
  • University of Zaragoza(萨拉戈萨大学)

机构由 AI 辅助整理,请以论文原文为准。

↑