发表机构
Thomas Jefferson High School for Science and Technology(托马斯·杰斐逊科学技术高中)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出StateSight基准及配套数据集StateSight-Steps,评估视觉语言模型的潜在空间状态重建能力,发现GPT-5.5、Claude Sonnet 5的表现均逊于人类基线,格式正确的响应可能掩盖空间结构恢复失败的问题。
AI 中文摘要
视觉语言模型正越来越多地用于多模态问答,然而很难单独评估其从单张图像重建潜在空间结构的能力。现有通用基准通常在同一评估中结合感知、光学字符识别、领域知识、语言先验和推理能力。我们推出StateSight,这是一个通过程序生成的基准,涵盖立方体对面推理、遮挡立方体塔计数以及四邻接连通分量计数三类任务。每个任务系列包含300个带确定神谕标签的单图像提示,采用精确匹配评分。使用API模型标识符gpt-5.5的OpenAI GPT-5.5在三项任务上的准确率分别为59.3%、33.3%和28.3%,Claude Sonnet 5的对应准确率为53.3%、18.7%和7.3%。所有直接运行的结果均无格式错误。针对60个样本的30人人类基线在每项任务上均超过了两个模型,平均准确率分别为80.8%、68.8%和64.3%。可见推导分析发现了图像状态重建和推理过程中的反复出现的错误。我们还推出配套数据集StateSight-Steps,包含900个交错的图像-文本示例和3600个确定的中间视觉状态。结果表明,格式有效的响应可能掩盖了恢复可验证视觉推理所需空间结构的失败。
英文摘要
Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct latent spatial structure from a single image remains difficult to isolate. Broad benchmarks often combine perception, optical character recognition, domain knowledge, linguistic priors, and reasoning in the same evaluation. We introduce StateSight, a procedurally generated benchmark for cube-net opposite-face reasoning, occluded cube-tower counting, and 4-neighbor connected-component counting. Each task family contains 300 single-image prompts with deterministic oracle labels and exact-match scoring. OpenAI GPT-5.5, using the API model identifier gpt-5.5, achieved 59.3%, 33.3%, and 28.3% accuracy across the three tasks, while Claude Sonnet 5 achieved 53.3%, 18.7%, and 7.3%. All final direct runs had zero format errors. A 30-participant human baseline on 60 items exceeded both models on every task, with mean accuracies of 80.8%, 68.8%, and 64.3%. Visible-derivation analysis identified recurring errors in image-state reconstruction and reasoning procedure. We also introduce StateSight-Steps, a companion dataset of 900 interleaved image-text examples and 3,600 deterministic intermediate visual states. The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.