发表机构
Preferred Networks, Inc.; The University of Tokyo(Preferred Networks 公司; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出T3lescope,一种无需逐场景优化的生成式表面重建方法,通过推理时由粗到细的级联固定分辨率生成器,在任意尺度下恢复高保真几何,并优于现有基线。
AI 中文摘要
我们在无需逐场景优化的情况下,从带位姿的多视图图像重建高保真3D场景网格,覆盖从单个物体到大型户外场景的不同尺度。逐场景优化方法在观测稀疏或表面具有光泽或透明特性时,缺乏所需的已学习3D先验。现有生成方法利用此类先验来补全稀疏观测区域的几何形状,但通常在有限空间范围内以固定分辨率运行,从而在空间覆盖与细节之间权衡。因此,重建大型场景通常需要将其划分为独立处理的、有重叠的局部区域,这使得维持全局几何一致性变得困难。为解决这些问题,我们提出T3lescope,它在推理时通过由粗到细的级联方式,跨场景尺度应用单一固定分辨率生成器。粗层级建立场景布局,更细层级在逐渐精细的空间单元内扰动并去噪从粗层级继承的几何形状,以恢复表面细节。该模型在多个尺度的单个单元上训练,并在所有层级间共享权重,因此训练时不固定层级结构,层级数量、单元尺度和单元位置均在推理时确定。在室内、室外和城市规模场景中,T3lescope优于前馈和生成式基线,匹配或超越逐场景优化,并恢复精细结构以及光泽和透明表面。这些结果表明,我们的方法能够泛化到多样场景、不同视图数量和图像分辨率。项目页面:此https URL
英文摘要
We reconstruct high-fidelity 3D scene meshes from posed multi-view images without per-scene optimization, across scales ranging from single objects to large outdoor scenes. Per-scene optimization methods lack the learned 3D prior needed when observations are sparse or surfaces are glossy or transparent. Existing generative methods leverage such priors to complete geometry in sparsely observed regions, but typically operate at a fixed resolution over a limited spatial extent, trading spatial coverage against detail. Reconstructing a large scene therefore often requires partitioning it into independently processed overlapping local regions, making it difficult to maintain global geometric consistency. To address these issues, we propose T3lescope, which applies a single fixed-resolution generator across scene scales in an inference-time coarse-to-fine cascade. A coarse level establishes the scene layout, and finer levels perturb and denoise geometry inherited from the coarser level within progressively finer spatial cells to recover surface detail. The model is trained on individual cells at multiple scales and shares its weights across all levels, so no hierarchy is fixed during training, and the number of levels, cell scales, and cell locations are determined at inference time. On indoor, outdoor, and city-scale scenes, T3lescope outperforms feed-forward and generative baselines, matches or surpasses per-scene optimization, and recovers fine structures as well as glossy and transparent surfaces. These results show that our method generalizes across diverse scenes, view counts, and image resolutions. Project page: https://pfnet-research.github.io/t3lescope/
Comments45 pages