arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22864cs.CVcs.AI

空间智商:通过分层能力测试解构空间智能

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

Patrick Rim, Tom Long, Ekta Prashnani, Ruth Rosenholtz, Ben Boudaoud, Peter Xenopoulos, Alex Wong, Joohwan Kim, Jae-Hyun Jung

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对多模态大语言模型空间推理能力不足问题,提出Spatial-IQ分层诊断框架,分解物体计数任务为子任务,生成数据集评估模型,发现模型存在问题,且用思维链监督和强化学习训练可提升模型表现。

中文摘要 AI 辅助

多模态大语言模型在视觉解释方面表现出色,但在人类能可靠解决的空间推理任务上却失败了。现有基准将这些模型作为黑箱评估,难以确定性能低下的根本原因。我们引入空间智商(Spatial-IQ),这是一个分层诊断框架,将堆叠3D结构中的物体计数分解为9个感知和认知子任务。利用NVIDIA Isaac Sim生成了约80000个带任务真值的数据集。评估了三种输出格式下的模型及人类基线。结果表明顶级模型常目标任务成功但底层子任务失败,模型在保留分层链程度上有差异。最后证明用思维链监督和可验证奖励的强化学习训练模型可提高子任务空间一致性和目标任务准确性。

英文摘要

Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Existing benchmarks evaluate these models as black boxes, limiting their ability to identify the underlying causes of lower performance: when a model fails a spatial reasoning task, it remains difficult to ascertain whether the hurdle is perceptual, such as recognizing object boundaries, or cognitive, such as reasoning about occlusion to infer hidden geometry. We introduce Spatial-IQ, a hierarchical diagnostic framework that decomposes object counting in stacked 3D structures into 9 perceptual and cognitive sub-tasks organized by the developmental stages of human spatial cognition, with mental rotation as an additional target probe. Using NVIDIA Isaac Sim, we procedurally generated a diverse dataset of roughly 80,000 stacked 3D structures with per-task ground truth. We evaluate models across three output formats (free-response text, multiple-choice images, and image editing) alongside a human baseline. The Spatial-IQ framework shows that top-performing models often succeed at the target task (object counting) without succeeding on the lower-level sub-tasks intended to support it, and that models differ in how much of these hierarchical chains they preserve, often revealing shortcut behavior that raw target-task accuracy alone would obscure. Finally, we demonstrate that training models with chain-of-thought (CoT) supervision over our hierarchical sub-tasks, combined with reinforcement learning with verifiable rewards, significantly improves both spatial consistency across sub-tasks and target-task accuracy, supporting the value of the proposed decomposition as both a diagnostic tool and a training signal.

发表机构

  • NVIDIA Research(英伟达研究院)
  • Yale University(耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑