arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.27033cs.ROcs.AIcs.CV

机器人中基于空间推理的视觉定位的多模态神经符号方法

A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics

  • Faculty of Information Technology, Monash University(信息技术学院,墨尔本大学)
  • Department of Electrical Engineering, Sharif University of Technology(电气工程系,谢赫大学)

机构由 AI 辅助整理,请以论文原文为准。

Simindokht Jahangard, Mehrzad Mohammadi, Abhinav Dhall, Hamid Rezatofighi

更新

AI总结:

针对VLMs在细粒度空间推理上的不足,提出整合全景图像与3D点云的神经符号框架,结合神经感知与符号推理显式建模空间逻辑关系,在JRDB-Reasoning数据集上展现卓越性能且保持轻量化设计。

AI中文摘要:

视觉推理,特别是空间推理,是一项具有挑战性的认知任务,要求理解复杂环境中的物体关系及其交互,尤其在机器人领域。现有的视觉语言模型(VLMs)擅长感知任务,但由于其隐式的、相关性驱动的推理及仅依赖图像,在细粒度空间推理上存在困难。我们提出了一种新颖的神经符号框架,整合全景图像与3D点云信息,将神经感知与符号推理相结合,以显式建模空间和逻辑关系。该框架包含用于检测实体和提取属性的感知模块,以及构建结构化场景图以支持精确、可解释查询的推理模块。在JRDB-Reasoning数据集上的评估表明,我们的方法在拥挤的人造环境中展现出卓越的性能和可靠性,同时保持了适合机器人与具身AI应用的轻量化设计。

英文摘要:

Visual reasoning, particularly spatial reasoning, is a challenging cognitive task that requires understanding object relationships and their interactions within complex environments, especially in robotics domain. Existing vision_language models (VLMs) excel at perception tasks but struggle with fine-grained spatial reasoning due to their implicit, correlation-driven reasoning and reliance solely on images. We propose a novel neuro_symbolic framework that integrates both panoramic-image and 3D point cloud information, combining neural perception with symbolic reasoning to explicitly model spatial and logical relationships. Our framework consists of a perception module for detecting entities and extracting attributes, and a reasoning module that constructs a structured scene graph to support precise, interpretable queries. Evaluated on the JRDB-Reasoning dataset, our approach demonstrates superior performance and reliability in crowded, human_built environments while maintaining a lightweight design suitable for robotics and embodied AI applications.

↑