发表机构
Electronics and Telecommunications Research Institute(韩国电子通信研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种无需训练的框架,通过局部上下文平铺生成与全局外观对齐,将3D生成先验扩展到大规模多视图场景网格,提升几何与外观一致性。
AI 中文摘要
预训练的3D生成模型能够产生精细的几何形状和外观,但主要设计用于有限空间范围内的以对象为中心的生成。最近的方法通过将大场景划分为较小的空间区域,并将预训练的3D生成先验应用于每个区域来解决这一局限性。然而,将平铺生成扩展到大型多视图场景时,保持局部几何连续性和全局外观一致性变得具有挑战性。我们提出了一种无需训练的大规模纹理网格生成框架,用于从多视图图像生成。我们的关键思想是将平铺生成扩展到具有更高空间细节的大型场景,同时在局部和全局上协调生成。我们引入了局部上下文平铺生成来改善相邻区域之间的几何连续性,以及全局外观对齐来减少远距离区域之间的外观差异。自适应场景分解进一步根据输入场景几何确定平铺数量。实验表明,与现有方法相比,我们的方法在几何和外观保真度上有所提升,同时能够实现大规模场景的细粒度生成。
英文摘要
Pretrained 3D generative models produce detailed geometry and appearance but are primarily designed for object-centric generation within a limited spatial extent. Recent approaches address this limitation by partitioning large scenes into smaller spatial regions and applying pretrained 3D generative priors to each region. However, scaling tiled generation to large multi-view scenes makes it challenging to maintain local geometric continuity and global appearance consistency. We present a training-free framework for large-scale textured mesh generation from multi-view images. Our key idea is to scale tiled generation to large scenes with increased spatial detail while coordinating generation both locally and globally. We introduce local context tiled generation to improve geometric continuity between neighboring regions and global appearance alignment to reduce appearance discrepancies across distant regions. An adaptive scene decomposition further determines the number of tiles according to the input scene geometry. Experiments demonstrate improved geometric and appearance fidelity over existing approaches while enabling fine-grained generation of large-scale scenes.