arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05242cs.LGcs.CV

将3D建模与空间推理解耦

Disentangling 3D Modeling from Spatial Reasoning

Haoze Sun, Jiequan Cui, Qingshan Xu, Richang Hong

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出解耦空间推理器(DiSR)框架,将3D感知与LLM推理解耦,在空间推理基准上取得竞争力性能,兼具可解释性、模块化性与计算效率,为空间智能提供替代范式。

中文摘要 AI 辅助

在本研究中,我们探索一种空间推理的替代范式,即通过显式将3D感知与推理解耦,而非通过大规模训练联合获取隐式3D感知与推理。我们的核心观察是,现代感知模型擅长估计连续的3D几何,而大型语言模型(LLMs)在组合式与符号式推理上尤为高效。受这些互补优势的启发,我们提出了解耦空间推理器(Disentangled Spatial Reasoner, DiSR),这是一个简单却有效的框架,它利用现成的专家感知模型将物理世界重构为结构化的3D证据,并通过LoRA微调大型语言模型,使其仅基于该显式几何证据执行推理。在未进行大规模3D视觉问答(VQA)训练或复杂工具使用策略的情况下,DiSR在流行的空间推理基准上取得了具有竞争力的性能。除了出色的性能,DiSR还具备更好的可解释性、模块化性和计算效率,证明感知与推理的显式分离是空间智能端到端建模的可扩展且有效的替代范式。

英文摘要

In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training. Our key observation is that modern perception models excel at estimating continuous 3D geometry, whereas large language models (LLMs) are particularly effective at compositional and symbolic reasoning. Motivated by these complementary strengths, we propose the Disentangled Spatial Reasoner (DiSR), a simple yet effective framework that reconstructs the physical world into structured 3D evidence using off-the-shelf expert perception models and fine-tunes an LLM with LoRA to perform reasoning solely over this explicit geometric evidence. Without large-scale 3D VQA training or complex tool-use policies, DiSR achieves competitive performance on popular spatial reasoning benchmarks. Beyond its strong performance, DiSR offers improved interpretability, modularity, and computational efficiency, demonstrating that explicit separation of perception and reasoning is a scalable and effective alternative paradigm to end-to-end modeling for spatial intelligence.

↑