arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于3D基础模型场景表示的零样本新视角深度合成

Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations

Denis M. Akola, David F. Fouhey

arXiv 2609.04174首次发表:更新:

发表机构

New York University; Tandon School of Engineering; Courant Institute of Mathematical Sciences(纽约大学; 坦登工程学院; 柯朗数学科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于3D基础模型的Z3D方法,通过对其内部表示做潜在扩散,实现零样本下多数据集新视角真实感深度图的预测。

AI 中文摘要

VGGT等3D基础模型(3DFMs)近期通过前馈Transformer预测丰富统一表示,推动了3D视觉的发展,其学习到的场景表示在多项3D视觉任务中表现出色。本文研究利用这些模型的内部表示从新视角推断场景的3D信息,假设为解决3D重建任务,这些模型需学习包含大量3D场景通用知识的表示。在证明可从3DFM内部表示解码隐藏表面后,本文提出方法Z3D,通过对3DFM表示进行潜在扩散来估计未见过视角的点云图。实验显示Z3D可在多个数据集上预测新视角的真实感深度图。

英文摘要

3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision by predicting rich unified representations with feed-foward transformers. The scene representations learned by these models enable strong performance on multiple 3D vision tasks. In this paper, we investigate using their internal representations to infer 3D in the scene from new views. Our hypothesis is that in order to solve the task of 3D reconstruction, these models need to learn a representation that includes a large amount of general knowledge about 3D scenes. After showing that it is possible to decode hidden surfaces from internal 3DFM representations, we propose a method, Z3D, that estimates pointmaps in unseen views by doing latent diffusion on 3DFM representation. We show that Z3D can predict realistic depth maps for new views across multiple datasets.

CommentsAccepted to the European Conference on Computer Vision (ECCV) 2026. Project page: https://akola-mbey-denis.github.io/Z3D-page/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑