arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpatioLM:面向视觉-语言模型的通用物理空间智能

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

Jing Wu, Jianhua Wu, Jiayi Guan, Jiahong Chen, Jinghui Lu, Hangjun Ye, Bingzhao Gao, Long Chen

arXiv 2608.01899首次发表:更新:

AI 中文总结

SpatioLM是一种参数高效的视觉-语言模型,通过内置空间视觉模块和伪深度、相机监督提升空间智能,在VSI-Bench获71.6分,还可迁移至具身操纵任务,同时保留通用能力。

AI 中文摘要

视觉-语言模型(VLMs)在常识推理任务上表现良好,但在视觉空间推理方面存在不足。现有多数解决方案会引入额外的3D先验输入或外部空间编码器,这会增加模型复杂度,且在空间微调后会削弱VLMs的通用能力。为此,我们提出参数高效的Spatio-vision Language Models(SpatioLM),该模型无需额外3D先验输入或第三方空间编码器即可提升空间智能。具体而言,我们设计了即插即用且非侵入式的空间视觉模块,以挖掘VLMs固有的空间知识;此外,我们创新性地利用伪深度和相机信息作为监督,引导模型学习物理一致性表征。大量实验表明,SpatioLM在空间感知与理解等多种任务上取得显著提升,同时有效限制了通用能力的下降。值得注意的是,该模型在VSI-Bench上取得了71.6的优异分数(是首个超过70的模型),在迁移到具身操纵任务时也达到了有竞争力的性能,代码可在https://github.com/spatio-lm获取。

英文摘要

Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \href{https://github.com/xiaomi-research/spatio-lm}{\faGithub~spatio-lm}.

Comments27 pages,13 figures,16 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑