发表机构
Brigham Young University(杨百翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究分析户外3DSG的挑战,以Terra 3DSG为案例,引入一致性指标,验证其用于物体检索的可行性,同时指出其在语义模式、可通行性等方面的开放问题。
AI 中文摘要
三维场景图(3DSG)已成为构建几何基础扎实、语义信息丰富、适用于高层机器人推理的通用分层地图的有前景方法。然而,3DSG在真实户外部署中的表现仍未被充分理解,尤其是结合开放集视觉语言模型(VLM)时。本领域报告分析了五个户外机器人数据集中大多数3DSG表示的通用组件,以表征复杂户外环境中出现的挑战。以近期提出的Terra 3DSG为案例研究,我们调查了五个不同数据集上的语义点嵌入、地点节点图导航、区域级理解和内存大小。我们还引入了新颖的一致性指标,以评估在同一环境的多次遍历中语义和结构图属性是否保持稳定。我们的分析显示,在所有测试数据集中,VLM点嵌入中异常值和多模式现象普遍存在,约30%的点的异常值比率高于0.1。我们展示了户外3DSG用于基于导航的物体检索的可行性,成功率接近70%,不过性能受可通行性故障和路径效率低下的限制,轨迹的次优路径效率平均约为66%。在复杂自然环境中,区域级理解仍然具有挑战性,平均F1分数较低,约为0.359。总体而言,我们的结果表明,户外3DSG能够保持紧凑(多千米轨迹的大小小于600MB)且相对一致的大规模环境表示,同时凸显了在处理多种语义模式、将可通行性纳入图结构以及改进高层区域理解方面的开放挑战。
英文摘要
Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounded, semantically informed, hierarchical general-purpose maps to support high-level robotic reasoning. However, the behavior of 3DSGs in real-world outdoor deployments remains poorly understood, particularly when combined with open-set vision-language models (VLMs). In this field report, we analyze the components common to most 3DSG representations across five outdoor robotic datasets to characterize challenges that arise in complex outdoor environments. Using the recently proposed Terra 3DSG as a case study, we investigate semantic point embeddings, place-node graph navigation, region-level understanding, and memory size across the five diverse datasets. We additionally introduce novel consistency metrics to evaluate whether semantic and structural graph properties remain stable across repeated traversals of the same environment. Our analysis reveals that outliers and multiple modes are common in VLM point embeddings across all tested datasets with outlier ratios above $0.1$ for around $30\%$ of points. We demonstrate the feasibility of outdoor 3DSGs for navigation-based object retrieval, achieving success rates near $70\%$, though performance is limited by traversability failures and inefficient routing, with trajectories averaging approximately $66\%$ suboptimal path efficiency. Region-level understanding remains challenging in complex natural environments, with low average F1 scores around $0.359$. Overall, our results show that outdoor 3DSGs can maintain compact (less than $600$MB for multi-kilometer trajectories) and relatively consistent large-scale environment representations, while highlighting open challenges in handling multiple semantic modes, incorporating traversability into graph structures, and improving higher-level region understanding.
CommentsThis work has been accepted for publication with the IEEE Transactions of Field Robotics Journal