发表机构
Universidade Federal de Mato Grosso (UFMT); Universidade Federal de Goiás (UFG); Advanced Knowledge Center for Immersive Technologies (AKCIT)(马托格罗索联邦大学; 戈亚斯联邦大学; 沉浸式技术高级知识中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SpatialChain基准,通过链重叠和场景图感知评判器审计VLM空间推理忠实度,发现高准确率下的捷径率,并验证SFT可提升忠实度。
AI 中文摘要
具备思考能力的视觉语言模型(VLM)在空间基准测试中报告了越来越高的准确率,然而最终答案的分数无法揭示正确的预测是反映了忠实的空间推理还是语言捷径。我们引入了SpatialChain,一个包含28,350个训练样本和899个测试样本的数据集,将面向空间的GQA问题与基于场景图的推理链配对,仅在生成的答案与符号真值匹配时保留,并采用双轴评估,将客观的链重叠指标与场景图感知的LLM评判器相结合,该评判器独立于最终答案对忠实度和完整性进行评分。应用于九个具备思考能力的VLM,该协议揭示了标准准确率无法显现的三个发现:(i)九个模型中有四个在VQA准确率达到≥79%的同时,表现出超过39%的捷径率,即正确答案的推理被评判器标记为不忠实;(ii)链质量显著预测了九个模型中七个的答案正确性,但两个例外(Claude Sonnet 4.6、InternVL3.5-8B)揭示了定性不同的失败模式——简洁输出与冗长装饰性推理——而仅凭基准准确率会混淆这些模式;(iii)在SpatialChain上进行SFT使Qwen3-VL-8B在领域内提高了+6.2个百分点,并将其捷径率降至22%,而对外部基准的风格特化效应促使采用重放增强训练作为缓解措施。忠实度评判器针对198个人工标注项进行了验证,其中评判器与人类的一致性达到了人类与人类的一致性水平,并针对来自不同提供商的第二个评判器进行了验证,该评判器保持了模型排名(ρ = 0.88)。数据、生成脚本和评估代码已在此https URL发布。
英文摘要
Thinking-enabled vision-language models (VLMs) report ever-higher accuracy on spatial benchmarks, yet final-answer scores cannot reveal whether a correct prediction reflects faithful spatial reasoning or a linguistic shortcut. We introduce SpatialChain, a dataset of 28,350 training and 899 test examples pairing spatially-oriented GQA questions with scene-graph-grounded reasoning chains, retained only when the generated answer matches the symbolic ground truth, and a two-axis evaluation combining objective chain-overlap metrics with a scene-graph-aware LLM judge that scores faithfulness and completeness independently of the final answer. Applied to nine thinking-enabled VLMs, the protocol surfaces three findings invisible to standard accuracy: (i) four of nine models achieve $\geq$79% VQA accuracy while exhibiting shortcut rates above 39%, i.e., correct answers whose reasoning the judge marks as unfaithful; (ii) chain quality significantly predicts answer correctness for seven of nine models, but the two exceptions (Claude Sonnet 4.6, InternVL3.5-8B) reveal qualitatively distinct failure modes, terse output vs. verbose-decorative reasoning, that benchmark accuracy alone conflates; (iii) SFT on SpatialChain improves Qwen3-VL-8B by +6.2 pp in-domain and reduces its shortcut rate to 22%, while a stylistic specialization effect on external benchmarks motivates replay-augmented training as mitigation. The faithfulness judge is validated against 198 human-annotated items, where judge-human agreement matches human-human agreement, and against a second judge from a different provider, which preserves the model ranking ($ρ$ = 0.88). Data, generation scripts, and evaluation code are released at https://github.com/spatialchain/SpatialChainBenchmark.
CommentsAccepted at the 2nd Workshop on Embodied Spatial Reasoning (ESR), NeurIPS 2026. 29 pages (8 main), 9 figures, 18 tables. Code and data: https://github.com/spatialchain/SpatialChainBenchmark