arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FPSGen:基于鸟瞰图支持的传输流的灵活点云场景生成

FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows

Wenzhe He, Meng Wang, JiaWei Qian, Jinfeng Xu, Ying Liu, Ruihui Li

arXiv 2607.26645首次发表:更新:

AI 中文总结

FPSGen是一种不依赖部分扫描结果的灵活点云场景生成框架,通过BEV先验构建点源与路径拉直传输方案,在SemanticKITTI补全和KITTI-360无条件生成任务上达到先进性能。

AI 中文摘要

现有的面向室外场景的基于点的生成方法主要关注激光雷达条件下的补全。训练时,通过扰动完整的真实场景构建含噪点云;推理时,则通过向重复的部分扫描结果添加噪声进行初始化。这种训练-推理不匹配继承了部分扫描结果的稀疏性和可见性偏差,导致远距离区域稀疏、遮挡区域几何结构不完整。此外,对部分扫描结果的依赖限制了激光雷达观测不可用或被布局线索替代时的生成能力。我们提出FPSGen,这是一种不依赖部分扫描结果独立构建点源的灵活框架。FPSGen首先从主动线索预测具有密度、高度和掩码通道的鸟瞰图(BEV)先验,随后对密度图进行采样以形成BEV支持的点源,支持无条件和有条件初始化。接着,采用师生近似最优传输方案,利用教师预测的端点学习诱导更直传输路径的速度场。通过将BEV点源构建与路径拉直传输相结合,FPSGen为无条件和灵活线索条件下的场景生成提供了统一框架。大量实验表明,FPSGen在SemanticKITTI补全任务上达到了先进的JSD和体素IoU性能,同时在单步点传输下保持了强性能;在KITTI-360无条件生成任务上,其还在对比方法中取得了最佳的覆盖率(COV)。

英文摘要

Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion. During training, noisy point clouds are constructed by perturbing complete ground-truth scenes, whereas during inference, they are initialized by adding noise to duplicated partial scans. This train-inference mismatch inherits the sparsity and visibility bias of partial scans, leading to sparse distant regions and incomplete geometry in occluded areas. Moreover, the reliance on partial scans restricts generation when LiDAR observations are unavailable or replaced by layout cues. We present FPSGen, a flexible framework that constructs point sources independently of partial scans. FPSGen first predicts a bird's-eye-view (BEV) prior with density, height, and mask channels from the active cues. The density map is then sampled to form a BEV-supported point source, enabling both unconditional and conditioned initialization. A teacher-student approximate optimal transport scheme then uses teacher-predicted endpoints to learn a velocity field that induces straighter transport paths. By integrating BEV point source construction with path-straightening transport, FPSGen provides a unified framework for unconditional and flexible cue-conditioned scene generation. Extensive experiments show that FPSGen achieves state-of-the-art JSD and voxel IoU performance on SemanticKITTI completion while maintaining strong performance with a single point transport step. On KITTI-360 unconditional generation, it also achieves the best Coverage (COV) among the compared methods.

Comments34 pages, 16 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑