SphMind:面向360度相机的鲁棒、免训练VLM空间推理
SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera
浏览论文内容
中文总结 AI 辅助
SphMind提出免训练即插即用框架,通过球谐空间图和推理时几何接地,解耦语义与几何,提升360度相机下MLLM的空间推理鲁棒性,在多个基准上显著优于基线。
中文摘要 AI 辅助
全方位或360度相机为具身智能体提供了周围环境的全景宽视场(FoV)视图,推动了多模态大语言模型(MLLMs)在全方位空间推理中的应用。然而,大多数MLLMs是在传统2D透视图像上训练的,难以应对球面几何引起的严重畸变和环绕不连续性问题。因此,在不重新训练的情况下使它们泛化到非欧几里得3D空间仍然具有挑战性。我们提出了SphMind,一个免训练、即插即用的框架,将语义感知与几何推理解耦。SphMind不需要MLLMs在内部学习球面几何,而是在外部处理几何的同时保留其语义能力。我们引入了基于球谐函数的空间图(SHSG),通过球面上的等变变换建模空间关系,以及推理时几何接地(IGG),一种模型无关的闭环优化过程,在推理期间将MLLM表示与球面几何约束对齐。在三个基准上的实验表明,SphMind在MP3D和Stanford2D-3D上的方向推理平均改进超过21.4%,在真实世界ODI-Bench上优于提示工程基线8.7%,并在全景旋转下将旋转不变性提高了5.9%,且无需额外训练或特定数据集调整。野外评估进一步表明,SphMind解决了基线视觉语言模型无法正确回答的方向推理查询。
英文摘要
Omnidirectional or 360 cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) view of their surroundings, motivating the use of Multi-modal Large Language Models (MLLMs) for omnidirectional spatial reasoning. However, most MLLMs are trained on conventional 2D perspective images and struggle with the severe distortions and wrap-around discontinuities induced by spherical geometry. Enabling them to generalize to non-Euclidean 3D spaces without retraining therefore remains challenging. We propose SphMind, a training-free, plug-and-play framework that decouples semantic perception from geometric reasoning. Rather than requiring MLLMs to learn spherical geometry internally, SphMind preserves their semantic capabilities while handling geometry externally. We introduce a Spherical Harmonics-based Spatial Graph (SHSG) that models spatial relationships through equivariant transformations on the sphere, together with Inference-Time Geometric Grounding (IGG), a model-agnostic closed-loop optimization process that aligns MLLM representations with spherical geometric constraints during inference. Experiments on three benchmarks show that SphMind achieves over 21.4% average improvement in directional reasoning on MP3D and Stanford2D-3D, outperforms prompt-engineering baselines by 8.7% on the real-world ODI-Bench, and improves rotational invariance by 5.9% under panorama rotations, without additional training or dataset-specific tuning. In-the-wild evaluations further show that SphMind resolves directional reasoning queries that baseline vision-language models fail to answer correctly.
发表机构
- NTU Singapore(新加坡南洋理工大学)
- IHPC, A*STAR(新加坡科技研究局高性能计算研究所)
机构由 AI 辅助整理,请以论文原文为准。