arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

看见、说出但不使用:多模态大语言模型中从可报告的空间事实到可用状态

Seeing, Saying, but Not Using: From Reportable Spatial Facts to Usable States in Multimodal Large Language Models

Jinchang Zhang, Guoyu Lu

arXiv 2610.02876首次发表:更新:

发表机构

Indiana University Bloomington(印第安纳大学布卢明顿分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SpaceConflict基准和操作状态监督(OSS),揭示多模态大模型能报告空间事实但未必在推理中使用,OSS通过监督状态及变换轨迹提升L3和L4层级的配对准确率。

AI 中文摘要

一个能正确报告空间事实的多模态大语言模型,未必会在后续推理中使用该事实。为研究这一区别,我们引入了SpaceConflict基准,包含23,196个输入,用于空间状态的构建与使用。在统一的“支持/矛盾/未知”判断界面下,该基准涵盖了局部事实绑定(L1)、关系组合(L2)、跨观测一致性(L3)以及变换下的状态判断(L4)。对同一世界提出直接状态查询、完整变换查询和显式初始状态查询,揭示了一种可用性-利用差距:模型能从视觉证据中恢复初始状态,但当该状态必须驱动变换时却失败。对于Qwen3.5-9B,100个正确恢复初始状态的序列中有50个在完整变换中失败,而显式提供状态则修复了全部50个;该差距随规模增大而缩小,但并未消除。因此,我们提出了操作状态监督(OSS),它监督任务相关的空间状态及其变换轨迹,并跨上下文对齐共享事实。OSS在匹配判断上的配对准确率提升最大的是L3和L4,即依赖于组织和使用状态的层级。因此,评估多模态空间推理不仅需要询问模型能否看见并陈述空间事实,还需询问该事实是否在后续计算中成为可用状态。

英文摘要

A multimodal large language model that correctly reports a spatial fact does not necessarily use that fact in subsequent reasoning. To study this distinction, we introduce \textsc{SpaceConflict}, a benchmark of 23{,}196 inputs for the construction and use of spatial state. Under a unified Supported/Contradictory/Unknown judgment interface, it covers local fact binding (L1), relational composition (L2), cross-observation consistency (L3), and state judgment under transformation (L4). Posing a direct-state query, a full-transformation query, and an explicit-initial-state query on the same world reveals an availability--utilization gap: models recover the initial state from visual evidence yet fail when that state must drive a transformation. For Qwen3.5-9B, 50 of 100 sequences with a correctly recovered initial state fail the full transformation, and supplying the state explicitly repairs all 50; the gap narrows with scale but does not close. We therefore propose Operational State Supervision (OSS), which supervises task-relevant spatial states and their transformation trajectories and aligns shared facts across contexts. OSS improves paired accuracy on matched judgments most on L3 and L4, the levels that depend on organizing and using state. Evaluating multimodal spatial reasoning thus requires asking not only whether a model can see and state a spatial fact, but whether that fact becomes a usable state in subsequent computation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑