arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.20837cs.CVcs.AI

ArchSIBench: 评估视觉-语言模型的建筑空间智能

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models

  • School of Architecture, Tsinghua University(清华大学建筑学院)

机构由 AI 辅助整理,请以论文原文为准。

Qirui Shen, Wenda Wang, Jiachen Lu, Zilong Huang, Jin Bai, Lei He, Hongxuan Chen, Weixin Huang

更新

AI总结:

本文提出ArchSIBench,一个基于建筑学、认知科学和心理学视角的建筑空间智能评估基准,通过17个细粒度子任务和3000个问题-答案对,评估多种VLMs在建筑空间感知、推理、导航、转换和配置方面的性能,发现大多数模型在空间转换和配置推理上仍与有建筑训练的人类评估者存在差距。

AI中文摘要:

建筑空间智能,即识别和推断建筑空间的能力,是机器人导航、具身交互和3D场景理解和生成等任务的基础。尽管已有大量研究评估了视觉-语言模型(VLMs)的基本空间技能,如相对方向、距离比较和物体计数,但这些任务仅涵盖空间认知的最基础层次,且忽略了更高层次的建筑空间认知,包括布局理解、通行模式和功能分区。在本文中,我们提出ArchSIBench,一个基于建筑学、认知科学和心理学视角的建筑空间智能评估基准。ArchSIBench涵盖五个核心维度:感知、推理、导航、转换和配置,包含17个细粒度子任务。通过专家的精心人工标注,我们构建了3,000个问题-答案对,以实现对建筑空间智能的全面评估。基于ArchSIBench,我们评估了各种VLMs,并发现大多数模型在建筑空间智能方面与人类基线有显著差异;此外,模型在能力维度上表现出显著的差异性。一些最先进的模型可以接近没有建筑训练的人类评估者水平。然而,与有建筑训练的人类评估者相比,仍存在明显差距,特别是在空间转换和配置推理方面。我们相信,ArchSIBench将为测量和提升VLMs的建筑空间智能提供重要的见解和系统资源。数据集和代码可在https://huggingface.co/datasets/ArchSIBench/ArchSIBench获取。

英文摘要:

Architectural spatial intelligence, the ability to recognize and infer architectural space, is fundamental to tasks such as robot navigation, embodied interaction, and 3D scene understanding and generation. Although extensive research has evaluated the basic spatial skills of Vision-Language Models (VLMs) such as relative orientation, distance comparison, and object counting, these tasks cover only the most elementary levels of spatial cognition and largely overlook higher-level cognition of architectural space, including layout understanding, circulation patterns, and functional zoning. In this work, we present ArchSIBench, a Benchmark for Architectural Spatial Intelligence based on the perspectives from architecture, cognitive science, and psychology. ArchSIBench covers five core dimensions: perception, reasoning, navigation, transformation, and configuration, comprising 17 fine-grained subtasks. Through careful manual annotation by experts with architectural backgrounds, we construct 3,000 question-answer pairs to enable comprehensive evaluation of architectural spatial intelligence. Based on ArchSIBench, we evaluate various VLMs and find that the architectural spatial intelligence of most models shows significant differences from human baselines; additionally, models exhibit substantial variability across capability dimensions. Some state-of-the-art models can approach the level of human evaluators without architectural training. However, a clear gap remains compared to human evaluators with architectural training, particularly in spatial transformation and configuration reasoning. We believe that ArchSIBench will provide important insights and systematic resources for measuring and advancing the architectural spatial intelligence of VLMs. The dataset and code are available at https://huggingface.co/datasets/ArchSIBench/ArchSIBench.

补充信息

↑