发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人员推出首个预测空间推理基准SpaceCast-Bench,评估21个视觉语言模型发现其性能远低于人类,微调后Qwen3-VL-4B在域外基准获显著提升。
AI 中文摘要
现有空间推理基准主要测试空间感知:读取输入中已可见的关系。然而现实世界的空间智能需要预测性空间推理:从观测中构建场景,预判干预如何改变场景,并推理未被观测到的结果。我们推出SpaceCast-Bench,这是首个直接且诊断性评估该能力的基准。该基准围绕观测-转换-推理框架构建,包含来自182个现实场景的3862个问题,涵盖16种任务类型,分为三个级别:静态感知、局部预测和全局预测,逐步要求场景理解、空间状态更新以及对未观测结果的关系推理。对21个模型的评估暴露出显著差距:最强模型仅达到58.0%,而人类表现为87.2%,空间专用模型则接近随机水平。控制分析进一步揭示,桥接视图对整合分布式观测至关重要,且明确的3D证据比生成的结果图像或视频更可靠地使模型受益。在我们通过程序生成的数据上进行微调,使Qwen3-VL-4B的性能从34.0%提升至65.7%,并在六个域外基准上实现了宏平均增益。
英文摘要
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluating 21 models exposes a stark gap: the strongest model reaches only 58.0% against 87.2% human performance, while spatially specialized models remain near random chance. Controlled analyses further reveal that bridge views are critical for integrating distributed observations, and that explicit 3D evidence benefits models more reliably than generated outcome images or videos. Fine-tuning on our programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7% with macro-average gains across six out-of-domain benchmarks.
CommentsCode: https://github.com/ZJU-REAL/SpaceCast-Bench Dataset: https://huggingface.co/datasets/hongxingli/SpaceCast-Bench