arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态空间推理的视觉信用审计

Visual Credit Audit for Multimodal Spatial Reasoning

Feixiang Liu, Qiang Qiu, Lanbo Sun, Nan Wei, Huawei Shen, Xueqi Cheng

arXiv 2607.27069首次发表:更新:

AI 中文总结

该研究提出视觉信用审计(VCA)方法,分解多模态空间推理基准的成功维度,发现部分模型决策正确但未获图像额外信用,验证了其在评估模型视觉关系响应上的有效性。

AI 中文摘要

封闭的是/否空间基准即使在图像提供的支持与无图像语境相差无几时也会奖励正确答案。在固定强制选择界面下,视觉信用审计(Visual Credit Audit, VCA)区分两个估计量:一是基准图像是否为模型声明的决策提供了比仅文本和空白对照更多的支持,二是模型是否对特定关系的视觉证据做出响应。第一项审计无需训练和标签,也不要求答案翻转;应用标签可得到依赖信用正确性(Dependence-Credited Correctness, D-CC),在正确项上,它等于相同对照下与黄金标准对齐的正增益,而预测对齐可将审计扩展至错误。在四个开放多模态大语言模型(MLLM)和两个空间基准上,12.73%-26.25%的决策正确但未获信用。匹配相同划分的图像置换使D-CC降低21.25-47.80个百分点,每对95%置信区间均大于零。固定像素关系对比及3×3证据源因子设计表明,空白对照无法识别关系响应。在受控的正确但未获信用的一致决策中,对关系反转的响应占81.57%-100.00%,合并后32.11%的决策会改变答案。对108个几何兼容编辑的独立审计结果提供了有界的自然图像对应性检验。VCA因此将基准成功分解为正确性、额外图像支持及关系一致响应三个维度。

英文摘要

Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 geometry-compatible edits provide a bounded natural-image correspondence check. VCA thereby decomposes benchmark success into correctness, additional image support, and relation-consistent response.

Comments20 pages, 2 figures. Code: https://github.com/SouthWinter/VCA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑