从生成视角探索空间智能
Exploring Spatial Intelligence from a Generative Perspective
- Zhejiang University(浙江大学)
- State Key Laboratory of CAD & CG(计算机辅助设计与图形学国家重点实验室)
- Ant Group(蚂蚁集团)
- Westlake University(西湖大学)
- Zhejiang University of Technology(浙江工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出GSI-Bench,首个评估生成空间智能的基准,通过空间 grounded 图像编辑量化空间合规性与编辑保真度,实验表明生成训练可提升空间推理能力。
AI中文摘要:
空间智能对多模态大语言模型至关重要,但现有基准主要从理解视角评估。本文探讨现代生成或统一多模态模型是否具备生成空间智能(GSI),即在图像生成中尊重和操控3D空间约束的能力,并提出GSI-Bench,首个通过空间 grounded 图像编辑量化GSI的基准。该基准包含GSI-Real和GSI-Syn两个互补组件,结合统一评估协议,实现可扩展、模型无关的空间合规性与编辑保真度评估。实验表明,在GSI-Syn上微调统一多模态模型在合成和真实任务中均取得显著提升,并显著提升下游空间理解能力,首次明确证明生成训练可有效增强空间推理能力,为多模态模型空间智能发展开辟新路径。
英文摘要:
Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective. We ask whether modern generative or unified multimodal models also possess generative spatial intelligence (GSI), the ability to respect and manipulate 3D spatial constraints during image generation, and whether such capability can be measured or improved. We introduce GSI-Bench, the first benchmark designed to quantify GSI through spatially grounded image editing. It consists of two complementary components: GSI-Real, a high-quality real-world dataset built via a 3D-prior-guided generation and filtering pipeline, and GSI-Syn, a large-scale synthetic benchmark with controllable spatial operations and fully automated labeling. Together with a unified evaluation protocol, GSI-Bench enables scalable, model-agnostic assessment of spatial compliance and editing fidelity. Experiments show that fine-tuning unified multimodal models on GSI-Syn yields substantial gains on both synthetic and real tasks and, strikingly, also improves downstream spatial understanding. This provides the first clear evidence that generative training can tangibly strengthen spatial reasoning, establishing a new pathway for advancing spatial intelligence in multimodal models.