发表机构
Nepal Applied Mathematics and Informatics Institute for research (NAAMII); INSAIT, Sofia University; State University of New York at Stony Brook(尼泊尔应用数学与信息学研究所; 索非亚大学INSAIT; 纽约州立大学石溪分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SceneBench是一个包含966个逼真3D场景的分层基准,通过人机协同生成超过18.3万标注节点,定义存在性、空间智能和QRA三类任务,揭示视觉-语言模型在层级与组合推理上的显著性能下降。
AI 中文摘要
视觉-语言模型在2D图像理解方面表现出色,但在3D空间推理方面仍然受限。当前基准的局限性阻碍了进展。首先,3D数据集通常依赖点云,这些点云捕获了几何信息,但丢弃了纹理、文字和材质等丰富的视觉特征。其次,标注将对象孤立处理,而忽略了现实世界中的层级组织(场景、房间、功能区、对象组)。第三,评估任务狭隘地聚焦于基础识别,而非多步空间推理。在此背景下,我们引入了SceneBench,这是一个包含966个通过高斯泼溅重建的逼真3D场景的基准,并带有跨越场景、房间、功能区、对象组和单个对象的密集层级语义标注。这些标注通过一个结合视觉-语言模型和约1,500小时人工迭代细化与验证的人机协同流程生成,产生了超过183K个带有文本描述和3D边界框的标注节点。基于这一表示,我们定义了三个评估任务:探测对象属性的基于存在性的问题、涵盖计数、大小比较、距离和方向关系的空间智能问题,以及需要在语义层级间进行多步推理的基于问题-推理-答案(QRA)三元组。使用最先进的视觉-语言模型进行的实验表明,虽然模型在基础识别任务上表现良好(例如,检测准确率高达85%),但在层级和组合推理上性能大幅下降(例如,计数准确率降至60%),这揭示了现有基准未捕获的局限性。SceneBench为在逼真的3D环境中开发和评估具备细粒度空间推理能力的模型提供了一个现实的测试平台。
英文摘要
Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat objects in isolation while ignoring real-world hierarchical organization (scenes, rooms, functional areas, object groups). Third, evaluation tasks focus narrowly on basic recognition rather than multi-step spatial reasoning. In this context, we introduce SceneBench, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects. These annotations are produced through a human-in-the-loop pipeline combining vision-language models with roughly 1,500 human-hours of iterative refinement and verification, producing over 183K annotated nodes with textual descriptions and 3D bounding boxes. Building on this representation, we define three evaluation tasks: Existence-Based Questions probing object attributes, Spatial Intelligence Questions covering counting, size comparison, distance, and directional relations, and Grounded Question-Reasoning-Answer (QRA) triplets requiring multi-step reasoning across semantic levels. Experiments with state-of-the-art vision-language models show that while models perform well on basic recognition tasks (e.g., up to 85% accuracy for detection), performance drops substantially on hierarchical and compositional reasoning (e.g., down to 60% for counting), revealing limitations not captured by existing benchmarks. SceneBench provides a realistic testbed for developing and evaluating models capable of fine-grained spatial reasoning in photorealistic 3D environments.