arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29793cs.CV

GridFlow:用于无缝城市级三维点云生成的结构化潜在流

GridFlow: Structured Latent Flow for Seamless City-Scale 3D Point Cloud Generation

Xinyu Wang, Muhammad Ibrahim, Atif Mansoor, Ajmal Mian

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出GridFlow框架,结合多模态条件生成城市级无缝三维点云,并构建City3D-MultiGen基准,实验显示其性能优于现有基线。

中文摘要 AI 辅助

从遥感数据生成逼真的三维城市环境对模拟、城市规划和混合现实具有重要意义,但现有点云生成方法仅限于单个对象或有界室内场景,无法应对城市级生成的规模、无缝拼接和部分可观测性挑战。我们提出GridFlow,这是一个多阶段框架,可基于卫星图像、语义分割图和数字表面模型(DSM)在城市级生成密集的彩色点云(每150米×150米图块含10^5个点)。Grid-Aligned VAE将每个图块编码为拓扑保留的潜在网格,其中令牌对应固定空间区域,支持空间一致的多模态条件和紧凑的潜在空间边缘一致性,该特性隐式对齐数千个边界点以实现跨图块的无缝生成。条件整流流模型从融合的多模态条件中合成几何潜在变量,面向方向的扩散着色器分别处理卫星可见的水平表面和被遮挡的垂直立面。为支持标准化评估,我们基于公开三维数据源构建了City3D-MultiGen基准,该基准包含来自墨尔本和伦敦的16.3万个密集标注图块,配有对齐的点云、卫星图像、语义图和高程数据。实验表明,GridFlow在所有几何指标上均优于适配的点云生成基线,且能在任意大的城市区域生成视觉连贯、边界无缝的彩色点云。我们的基准详情可在该https URL获取。

英文摘要

Generating realistic 3D city environments from remote sensing data is important for simulation, urban planning, and mixed reality, yet existing point cloud generation methods are limited to single objects or bounded indoor scenes and cannot handle the scale, seamless tiling, and partial observability challenges of city-scale generation. We present \ours{}, a multi-stage framework that generates dense, colored point clouds ($10^5$ points per $150\text{m}{\times}150\text{m}$ tile) at city scale, conditioned on satellite imagery, semantic segmentation maps, and digital surface models (DSM). A \emph{Grid-Aligned VAE} encodes each tile into a topology-preserving latent grid where tokens correspond to fixed spatial regions, enabling spatially coherent multi-modal conditioning and compact latent-space edge consistency that implicitly aligns thousands of boundary points for seamless cross-tile generation. A conditional rectified flow model synthesizes geometry latents from the fused multi-modal conditions, and an orientation-aware diffusion colorizer separately handles satellite-visible horizontal surfaces and occluded vertical façades. To support standardized evaluation, we build on public 3D data sources to introduce \emph{City3D-MultiGen}, a benchmark of $163$K densely annotated tiles from Melbourne and London with aligned point clouds, satellite images, semantic maps, and elevation data. Experiments show that \ours{} outperforms adapted point cloud generation baselines across all geometry metrics and produces visually coherent colored point clouds with seamless boundaries over arbitrarily large urban extents. Our benchmark details are available at https://huggingface.co/datasets/e32/City3D-MultiGen

发表机构

  • The University of Western Australia(西澳大学)

机构由 AI 辅助整理,请以论文原文为准。

↑