发表机构
Worcester Polytechnic Institute(伍斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人在非结构化环境中难以区分空间与语义相似容器的问题,提出基于空间图的LOCUS方法,通过GNN更新CLIP嵌入联合推理,在仿真中优于多种基线,并验证了物理系统的鲁棒性。
AI 中文摘要
随着机器人越来越多地被部署在非结构化的真实世界环境中,对物体之间复杂的空间和语义关系进行推理的能力,仍然是实现鲁棒且可泛化的操作与导航的一项基本挑战。例如,决定在哪里搜索一个不在视线范围内的物体,既取决于该物体在物理上可能的位置,也取决于其在语义上可能的位置。诸如使用CLIP的余弦相似度等流行方法,难以区分在空间和语义上相似且可能包含目标物体的容器。我们提出了基于空间图的地标导向容器判别方法(LOCUS)。为了联合推理空间信息和语义,我们训练了一个图神经网络(GNN),通过基于邻近性在节点间传递嵌入信息,在完整的环境场景图上更新CLIP嵌入。由融合了家庭本体论的语义知识增强的CLIP嵌入,提供了稳健的语义信号。我们在仿真中,基于模拟器元数据提供的无噪声场景图进行评估,从而使该方法与检测无关。在仿真中,我们的方法在大多数房间类别和配置中优于随机方法、CLIP、Tidybot以及一个LLM规划器。为了展示我们的方法对场景图节点位置和标签噪声的鲁棒性,我们演示了一个物理移动操作器在探索流程中端到端运行,包括来自Detic的场景图标签。
英文摘要
As robots are increasingly deployed in unstructured, real-world environments, the ability to reason about complex spatial and semantic relationships among objects remains a fundamental challenge in enabling robust and generalizable manipulation and navigation. For example, deciding where to search for an object that is not in plain sight depends on where it is physically plausible as well as semantically likely. Popular methods such as cosine similarity with CLIP struggle to disambiguate between spatially and semantically similar objects that could contain a target object. We propose Landmark-Oriented Container Discrimination Using Spatial Graphs (LOCUS). To jointly reason about spatial information and semantics, we train a GNN to update CLIP embeddings on a full environment scene graph by passing embedding information between nodes based on proximity. CLIP embeddings, augmented by semantic knowledge from a fusion of household ontologies provide a robust semantic signal. We evaluate in simulation on a noiseless scene graph from simulator metadata, making the approach detection-agnostic. Our approach outperforms random, CLIP, Tidybot, and an LLM planner in the majority of room classes and configurations in simulation. To show our approach is robust to scene graph node placement and label noise, we demonstrate a physical mobile manipulator running in an exploration pipeline start-to-finish including scene graph labels from Detic.