GS-Voxel:用于大规模3DGS生成的免拟合结构化隐变量
GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation
浏览论文内容
中文总结 AI 辅助
GS-Voxel是免拟合结构化隐变量框架,可将预优化3DGS重建转为稀疏活跃体素,通过GS专用VAE编码为稀疏3D隐变量,结合图像条件流模型实现大规模航空3DGS场景生成。
中文摘要 AI 辅助
许多可扩展的隐变量3D生成器基于结构化张量运行,而预优化的3D高斯溅射(3DGS)重建结果是无序的、空间不规则的,且基元数量差异很大。我们提出GS-Voxel,一种免拟合的结构化隐变量框架,并将其用于大规模航空3D高斯场景生成任务。GS-Voxel可确定性地将兼容的预优化3DGS重建结果转换为稀疏活跃体素,无需额外的逐场景优化,同时保留所选基元的亚体素位置和渲染属性。一种专门针对GS的因子化变分自编码器(VAE)随后将体素几何和局部高斯属性分别编码为稀疏3D隐变量,其大小随占用体素数量增长,而非受限于固定的全场景基元数量。我们在GS-Voxel隐变量空间中训练图像条件流模型,以生成航空3DGS场景。GS-Voxel实现的一项关键应用是大面积场景生成:重叠感知的分块推理可将合成范围扩展至单个训练裁剪块之外,该过程以卫星视图图像为条件。我们的结果表明,GS-Voxel为预优化的航空3DGS重建提供了结构化隐变量,其隐变量容量随占用体素数量增长。
英文摘要
Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimized 3DGS reconstruction into sparse active voxels without additional per-scene optimization, retaining the sub-voxel positions and rendering attributes of the selected primitives. A GS-specific factorized VAE then separately encodes voxel geometry and local Gaussian attributes into sparse 3D latents whose size grows with the number of occupied voxels rather than being limited by a fixed scene-wide primitive count. We train image-conditioned flow models in the GS-Voxel latent space to generate aerial 3DGS scenes. A key application enabled by GS-Voxel is large-area scene generation: overlap-aware tiled inference extends synthesis beyond a single training crop conditioned on satellite-view images. Our results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels.
发表机构
- Zhejiang University(浙江大学)
- Peking University(北京大学)
- Amap(高德地图)
- Alibaba(阿里巴巴)
机构由 AI 辅助整理,请以论文原文为准。