arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03892cs.CVcs.AIcs.RO

GraFT:一种基于3D场景图的多模态大语言模型空间推理无训练框架

GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs

Junqing Du, Fernando Ropero, Erkin Turkoz, Yanfeng Zhang, Lu Liu

首次发表
浏览论文内容

中文总结 AI 辅助

GraFT是一种无训练框架,通过3D场景图为多模态大语言模型提供三类空间推理能力,在ScanQA、VSI-Bench数据集上显著提升了模型空间推理性能。

中文摘要 AI 辅助

3D空间推理是理解和作用于物理世界的基础,但当前多模态大语言模型(MLLMs)的这类推理仍不可靠。这些模型在精确几何测量、自我中心与非自我中心视角转换、细粒度外观定位方面表现不佳。最常见的解决方法是在大规模精心整理的空间推理数据集上微调模型,或附加专用的3D几何编码器,这些方法通常将解决方案与高成本监督及特定主干网络绑定。我们提出GraFT,一种无训练框架,通过紧凑、易维护的3D场景图(3DSG)提供缺失的3D结构。基于该3DSG,GraFT提供三类空间推理能力:(1)通过符号工具实现确定性几何;(2)通过鸟瞰图(BEV)渲染实现非自我中心布局;(3)通过任务相关的自我中心帧实现视觉属性定位。在ScanQA数据集上,GraFT在相同主干网络的基线模型上提升了所有指标,将CIDEr提升了27%;在VSI-Bench上,GraFT使冻结的MLLMs性能提升高达65%,超过所有专有及通用开源基线模型,以及多个知名微调空间模型。

英文摘要

3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in current multimodal large language models (MLLMs). These models falter at precise geometric measurement, at transforming between egocentric and allocentric viewpoints, and at grounding fine-grained appearance. The most common remedies fine-tune the model on large-scale curated spatial-reasoning datasets or attach dedicated encoders for 3D geometry, which typically couples the solution to costly supervision and a specific backbone. We instead introduce GraFT, a training-free framework that supplies the missing 3D structure through a compact, easily maintained 3D scene graph (3DSG). From this 3DSG, GraFT provides three spatial reasoning capabilities: (1) deterministic geometry through symbolic tools, (2) allocentric layout through a bird's-eye-view (BEV) rendering, and (3) visual-attribute grounding through task-relevant egocentric frames. On ScanQA, GraFT improves every metric over the same-backbone baseline, raising CIDEr by 27%. On VSI-Bench, GraFT improves frozen MLLMs by up to 65%, surpassing every proprietary and general-purpose open-source baseline, and several prominent fine-tuned spatial models.

发表机构

  • Riemann Lab, Huawei Technologies(华为技术有限公司黎曼实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑