arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Imagine3D-LLM:在回答前教会多模态大语言模型想象3D场景

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han, Mungyeom Kim, Minkyeong Jeon, Heeseong Shin, Wonjun Moon, Federico Tombari, Daniel Barath, Marc Pollefeys, Seungryong Kim, Sunghwan Hong

arXiv 2609.38177首次发表:更新:

发表机构

KAIST AI; ETH Zürich; Google; TUM; ETH AI Center(韩国科学技术院人工智能学院; 苏黎世联邦理工学院; 谷歌; 慕尼黑工业大学; 苏黎世联邦理工学院人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

受人类空间推理启发,提出Imagine3D-LLM,通过可学习摘要标记解码为紧凑3D高斯泼溅表示并联合训练,使MLLM在回答前想象3D场景,显著提升空间推理与3D理解性能。

AI 中文摘要

从多视角图像推理3D世界仍然是多模态大语言模型(MLLMs)面临的基本挑战。虽然现代MLLMs能有效处理单张图像输入,但它们难以将跨视角的证据整合为连贯的3D理解。越来越多的研究工作试图通过向MLLMs注入3D感知来弥合这一差距,要么增强细粒度的像素级跨视角对应关系,要么融合来自3D几何基础模型的特征,但与人类推理能力之间仍存在显著差距。在本工作中,我们重新审视人类空间推理过程,该过程表明,人类并非依赖细粒度的几何线索,而是粗略地识别跨视角的常见物体,推断视角之间的相对几何关系,并组装出场景的粗略3D布局。受此过程启发,我们提出了Imagine3D-LLM,一种学习组装类似的紧凑3D场景表示并基于该表示来生成答案的MLLM。具体而言,我们在图像标记之后附加一小部分可学习的摘要标记,将其解码为紧凑的3D高斯泼溅表示,并通过光度重建损失进行监督,同时与标准的下一标记预测目标联合训练。值得注意的是,尽管只有摘要标记直接接收重建监督,该目标也增强了LLM底层图像特征中的跨帧对应关系,这表明学习重建会将3D感知信号传播到整个模型中。因此,Imagine3D-LLM在多个空间推理和3D理解基准上持续优于先前的方法,表明想象场景可能比被告知像素级几何更有效。

英文摘要

Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.

CommentsNeurIPS 2026; Project Page: https://cvlab-kaist.github.io/Imagine3D-LLM

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑