发表机构
KAIST; AITRICS(韩国科学技术院; AITRICS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LVLMs空间智能不足,提出基于合成积木堆叠任务的新范式,构建SpatialBlock-15k数据集,通过训练显著提升模型空间推理能力并泛化至真实任务。
AI 中文摘要
大型视觉-语言模型(LVLMs)在各种视觉任务上取得了强劲性能,但它们在从2D图像中重建和推理场景3D结构的能力——即空间智能——仍然有限。现有方法尝试通过使用需要密集几何标注的真实场景空间问答数据集来弥补这一差距。然而,由于依赖外部感知模块,构建此类标签成本高昂、耗时且常常带有噪声。在这项工作中,我们提出了一种受人类认知发展启发的新范式:通过结构化的积木操作任务学习基础空间技能。我们引入了SpatialBlock-15k,一个包含15,000个积木堆叠问题的合成数据集,涵盖3D到2D投影、视角变换和结构组合。该数据集进一步融入了受控的颜色调制作为视觉线索,以鼓励在视觉复杂条件下基于锚点的推理。实验表明,在我们的数据集上通过直接回答或基于推理的预测训练的LVLMs显著优于基线,并能泛化到真实世界的空间任务,尽管该数据集具有合成性和紧凑性。代码和数据可在该https URL获取。
英文摘要
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images -- referred to as spatial intelligence -- remains limited. Existing approaches attempt to address this gap by using real-scene spatial question answering datasets that require dense geometric annotations. However, constructing such labels is costly, time-consuming, and often noisy due to reliance on external perception modules. In this work, we propose a novel paradigm inspired by human cognitive development: learning foundational spatial skills through structured block-manipulation tasks. We introduce SpatialBlock-15k, a synthetic dataset of 15,000 block-stacking problems covering 3D-to-2D projection, viewpoint transformation, and structural combination. The dataset further incorporates controlled color modulation as visual cues to encourage anchor-based reasoning in visually complex conditions. Experiments demonstrate that LVLMs trained on our dataset through either direct answering or reasoning-based prediction significantly outperform baselines and generalize to real-world spatial tasks, despite the dataset's synthetic and compact nature. Code and data are available at https://github.com/rsoohyun/SpatialBlock.