arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

追踪、验证与修正:一种用于多模态大语言模型空间推理的无训练框架

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs

Yang Yang, Jiawei Chen, Tairan Chen, Zhaoxia Yin

arXiv 2608.04759首次发表:更新:

发表机构

East China Normal University; Zhongguancun Academy; Stevens Institute of Technology(华东师范大学; 中关村学院; 史蒂文斯理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型空间推理的错误传播问题,提出无训练的追踪-验证-修正框架,构建空间证据图并评估证据可靠性,在15种模型-数据集设置下平均准确率达68.94%,优于基线。

AI 中文摘要

尽管多模态大语言模型(Multimodal Large Language Models, MLLMs)已取得显著进展,但其空间推理仍可能产生与输入图像不一致的中间判断,致使错误在推理链中传播并影响最终答案。现有方法主要通过训练或补充空间信息来提升空间推理能力,未考虑推理过程本身是否忠实于模型输入。本研究表明,不忠实的推理链会显著降低最终答案的准确性。为解决该问题,我们提出一种模块化的无训练框架,用于空间推理的验证与修正。该框架构建空间证据图(Spatial Evidence Graph, SEG),将从思维链(Chain-of-Thought)推理中提取的原子空间证据与视觉实体、空间关系、源步骤及视觉证据相关联;空间证据可靠性评估(Spatial Evidence Reliability Assessment, SERA)则基于对象存在性、定位和几何测量评估视觉证据的可靠性。随后,框架识别出最早与可靠视觉证据相矛盾的空间证据单元,并指导原多模态大语言模型修正后续推理及最终答案。在15种模型-数据集设置下,我们的方法达到68.94%的平均准确率,较对比基线平均高出8.55个百分点,代码将开源。

英文摘要

Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inconsistent with the input image, allowing errors to propagate through the reasoning chain and affect the final answer. Existing methods mainly improve spatial reasoning through training or additional spatial information, without considering whether the reasoning process itself is faithful to the model input. Our study shows that unfaithful reasoning chains significantly reduce final-answer accuracy. To address this issue, we propose a modular and training-free framework for spatial reasoning verification and correction. The framework constructs a Spatial Evidence Graph (SEG), which associates atomic spatial evidence extracted from Chain-of-Thought reasoning with visual entities, spatial relations, source steps, and visual evidence. Spatial Evidence Reliability Assessment (SERA) evaluates the reliability of visual evidence based on object existence, localization, and geometric measurements. The framework then identifies the earliest spatial evidence unit contradicted by reliable visual evidence and guides the original MLLM to revise the subsequent reasoning and final answer. Across 15 model-dataset settings, our method achieves an average accuracy of 68.94%, outperforming the compared baselines by 8.55 percentage points on average. Our code will be open-sourced.

Comments19 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑