发表机构
KAIST AI; ETH Zürich; Google; TUM; ETH AI Center(韩国科学技术院人工智能学院; 苏黎世联邦理工学院; 谷歌; 慕尼黑工业大学; 苏黎世联邦理工学院人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
受人类空间推理启发,提出Imagine3D-LLM,通过可学习摘要标记解码为紧凑3D高斯泼溅表示并联合训练,使MLLM在回答前想象3D场景,显著提升空间推理与3D理解性能。
AI 中文摘要
从多视角图像推理3D世界仍然是多模态大语言模型(MLLMs)面临的基本挑战。虽然现代MLLMs能有效处理单张图像输入,但它们难以将跨视角的证据整合为连贯的3D理解。越来越多的研究工作试图通过向MLLMs注入3D感知来弥合这一差距,要么增强细粒度的像素级跨视角对应关系,要么融合来自3D几何基础模型的特征,但与人类推理能力之间仍存在显著差距。在本工作中,我们重新审视人类空间推理过程,该过程表明,人类并非依赖细粒度的几何线索,而是粗略地识别跨视角的常见物体,推断视角之间的相对几何关系,并组装出场景的粗略3D布局。受此过程启发,我们提出了Imagine3D-LLM,一种学习组装类似的紧凑3D场景表示并基于该表示来生成答案的MLLM。具体而言,我们在图像标记之后附加一小部分可学习的摘要标记,将其解码为紧凑的3D高斯泼溅表示,并通过光度重建损失进行监督,同时与标准的下一标记预测目标联合训练。值得注意的是,尽管只有摘要标记直接接收重建监督,该目标也增强了LLM底层图像特征中的跨帧对应关系,这表明学习重建会将3D感知信号传播到整个模型中。因此,Imagine3D-LLM在多个空间推理和3D理解基准上持续优于先前的方法,表明想象场景可能比被告知像素级几何更有效。
英文摘要
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.
CommentsNeurIPS 2026; Project Page: https://cvlab-kaist.github.io/Imagine3D-LLM