arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

空间思维链:面向视觉语言模型的模态无关空间定位

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian, Richard Shi, Jinjun Shan, Amir Rasouli, Dongfeng Bai

arXiv 2608.10278首次发表:更新:

AI 中文总结

该研究提出轻量架构无关框架Space Tokens,将空间信息蒸馏为连续token融入VLMs的思维链,在VSI-Bench上提升两款模型性能,在尺寸估计任务达SOTA,实现高效空间推理。

AI 中文摘要

空间理解是具身智能的基础,支撑着机器人操作、具身导航、自动驾驶等应用。尽管近期视觉语言模型(VLMs)在空间推理基准上取得了令人瞩目的性能,但当前最先进的方法通常在推理阶段依赖额外的空间编码器或架构修改,这增加了计算成本。我们提出Space Tokens,这是一种轻量、架构无关的框架,可为VLMs配备显式的连续空间表示,无需额外的推理模块。通过将场景级3D几何和以对象为中心的空间属性蒸馏为连续潜变量token,我们的方法使这些模态能够直接融入思维链推理过程,从而提升VLM的空间推理能力。同时,学习到的表示可被显式解码,以验证其是否编码了有意义的几何信息,而统一的token接口仍可扩展至其他模态。在VSI-Bench上的实验显示,该方法使Qwen3-VL-8B提升了4.3%,使SenseNova-SI-1.3提升了1.3%,并在对象尺寸估计(79.2%)和房间尺寸估计(75.7%)上达到了最先进的性能。这些结果表明,连续空间token为将几何推理集成到大型视觉语言模型中提供了一种有效、可解释且计算高效的机制。

英文摘要

Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent vision-language models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, state-of-the-art approaches typically rely on additional spatial encoders or architectural modifications during inference, increasing computational cost. We introduce Space Tokens, a lightweight, architecture-agnostic framework that equips VLMs with explicit continuous spatial representations without requiring additional inference-time modules. By distilling scene-level 3D geometry and object-centric spatial attributes into continuous latent tokens, our method enables these modalities to be directly incorporated into a chain-of-thought reasoning process, thereby improving the VLM's spatial reasoning capabilities. At the same time, the learned representations can be explicitly decoded to verify that they encode meaningful geometric information, while the unified token interface remains extensible to additional modalities. Experiments on VSI-Bench improve Qwen3-VL-8B by 4.3% and SenseNova-SI-1.3 by 1.3%, while achieving state-of-the-art performance on object size (79.2%) and room size estimation (75.7%). These results demonstrate that continuous spatial tokens provide an effective, interpretable, and computationally efficient mechanism for integrating geometric reasoning into large vision-language models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑