机器人中基于空间推理的视觉定位的多模态神经符号方法
A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
- Faculty of Information Technology, Monash University(信息技术学院,墨尔本大学)
- Department of Electrical Engineering, Sharif University of Technology(电气工程系,谢赫大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对VLMs在细粒度空间推理上的不足,提出整合全景图像与3D点云的神经符号框架,结合神经感知与符号推理显式建模空间逻辑关系,在JRDB-Reasoning数据集上展现卓越性能且保持轻量化设计。
AI中文摘要:
视觉推理,特别是空间推理,是一项具有挑战性的认知任务,要求理解复杂环境中的物体关系及其交互,尤其在机器人领域。现有的视觉语言模型(VLMs)擅长感知任务,但由于其隐式的、相关性驱动的推理及仅依赖图像,在细粒度空间推理上存在困难。我们提出了一种新颖的神经符号框架,整合全景图像与3D点云信息,将神经感知与符号推理相结合,以显式建模空间和逻辑关系。该框架包含用于检测实体和提取属性的感知模块,以及构建结构化场景图以支持精确、可解释查询的推理模块。在JRDB-Reasoning数据集上的评估表明,我们的方法在拥挤的人造环境中展现出卓越的性能和可靠性,同时保持了适合机器人与具身AI应用的轻量化设计。
英文摘要:
Visual reasoning, particularly spatial reasoning, is a challenging cognitive task that requires understanding object relationships and their interactions within complex environments, especially in robotics domain. Existing vision_language models (VLMs) excel at perception tasks but struggle with fine-grained spatial reasoning due to their implicit, correlation-driven reasoning and reliance solely on images. We propose a novel neuro_symbolic framework that integrates both panoramic-image and 3D point cloud information, combining neural perception with symbolic reasoning to explicitly model spatial and logical relationships. Our framework consists of a perception module for detecting entities and extracting attributes, and a reasoning module that constructs a structured scene graph to support precise, interpretable queries. Evaluated on the JRDB-Reasoning dataset, our approach demonstrates superior performance and reliability in crowded, human_built environments while maintaining a lightweight design suitable for robotics and embodied AI applications.