发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究图像生成模型空间认知评估问题,提出ProVisE框架及SpatialGen - Bench基准,通过统一设置评估相关模型,发现图像生成模型在像素空间直接外化答案时有竞争力,文本输出语言模型在组合空间推理中有优势,揭示了两者互补优势并建立测试平台。
AI 中文摘要
空间智能对于智能体从静态语义理解转向与物理世界交互至关重要。许多空间任务基于连续视觉场景,用指向、标记或绘图来表达位置等更自然。现有空间推理基准通常需要坐标等,与图像生成模型存在答案接口不匹配。我们提出ProVisE框架,它能从图像生成模型引出视觉答案并解析为结构化预测,还包含一个构建器。我们还引入SpatialGen - Bench基准。评估结果表明图像生成模型在直接于像素空间外化空间答案时有竞争力,而文本输出语言模型在组合空间推理中有明显优势。这些发现揭示了像素空间表达和基于文本推理的互补优势,并建立了用于研究图像生成模型中空间认知的度量兼容测试平台。
英文摘要
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.
Comments36 pages, 14 figures. Project page: https://zju-omniai.github.io/ProVisE/