arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33616cs.CVcs.AI

SpatialSpeak:面向空间链式思维推理的问答原生重建与局部及全局上下文

SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning

Yang Cao, Jiaxin Zhang, Dave Zhenyu Chen, Yingji Zhong, Ruiyuan Gao, Lanqing Hong, Dan Xu

首次发表
浏览论文内容

中文总结 AI 辅助

SpatialSpeak通过两阶段框架,以问答原生重建预训练结合空间链式思维推理,提升视觉语言模型的多视角空间推理能力,并在多个基准上取得最优结果。

中文摘要 AI 辅助

视觉语言模型(VLM)可受益于几何先验以进行多视角空间推理,然而仅基于答案的训练并不能直接监督中间几何估计及其在推导定量空间答案中的使用。我们假设,当VLM首先通过多视角重建联合学习互补的局部几何与全局场景上下文时,空间链式思维(CoT)监督将更加有效。我们提出SpatialSpeak,一个将问答原生重建预训练与空间CoT学习相连接的两阶段框架。在阶段I中,问答原生重建预训练(QA-RP)结合标记点3D查询以获取精细的局部几何,以及对象中心查询以获取跨视角的全局场景上下文。这两个任务均被表述为基于文本的问答,使得几何估计与后续推理能够共享同一自回归输出接口。在阶段II中,带视觉补偿的空间CoT(CoT-VC)训练模型表达与问题相关的几何估计并利用其推导答案,并辅以可靠性评估与视觉补偿,在需要时支持答案细化。在ReVSI上,QA-RP将CoT-VC带来的增益从2.6分提升至6.9分,且消融实验表明局部与全局重建监督均有益。SpatialSpeak在ReVSI、VSI-Bench和SPAR-Bench上取得了最先进的结果,其中ReVSI得分为62.8,超过最强对比基线8.7分。

英文摘要

Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction. We introduce SpatialSpeak, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed. On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.

发表机构

  • Hong Kong University of Science and Technology(香港科技大学)
  • Harbin Institute of Technology(哈尔滨工业大学)
  • Huawei Noah’s Ark Lab(华为诺亚方舟实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑