arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式语义场景补全

Generative Semantic Scene Completion

Shi Chen, Weifeng Ge

arXiv 2608.26737首次发表:更新:

发表机构

College of Computer Science and Artificial Intelligence, Fudan University; Robotics Institute, Carnegie Mellon University(复旦大学计算机科学与人工智能学院; 卡内基梅隆大学机器人研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对户外激光雷达语义场景补全的类别不平衡问题,提出生成式语义场景补全框架,通过PS³、SGSC、S²D²方法提升性能,在SemanticKITTI测试集取得当前最佳单扫描单样本结果。

AI 中文摘要

户外激光雷达语义场景补全(SSC)需从仅观测到目标体积1%的扫描数据中恢复密集的语义体素网格,且面临超过7000倍的类别不平衡问题。我们将SSC重新定义为生成式语义场景补全(GSSC),采用兼具三重角色的离散扩散公式:第一,配对稀疏-密集场景合成(PS³)生成匹配的稀疏激光雷达观测数据及其密集语义补全结果,从源头上解决长尾问题,并构建了用于训练的PS³-SemanticKITTI数据集,同时结合SemanticKITTI数据集使用;第二,语义引导的生成式场景补全(SGSC)通过多项式离散扩散从噪声中生成场景,以鸟瞰视角语义图和稀疏3D特征流作为条件,基于稀疏扫描进行生成;第三,同一框架可通过一次流匹配步骤优化现有补全结果:结构化源离散扩散(S²D²)。S²D²无需对基础模型进行重新训练或测试时适配,即可提升SGSC自身输出及所有测试的外部SSC基础模型的平均交并比(mIoU)。在最强基础模型上,不使用测试时增强的单步操作在SemanticKITTI隐藏测试集上达到38.8%的mIoU,据我们所知,这是该排行榜上符合因果、单扫描、单样本要求的最佳结果,比相同限制条件下之前发表的最佳分数高出2.1个百分点;使用八视图测试时增强的四步修正达到39.2%,超出该限制条件。

英文摘要

Outdoor LiDAR semantic scene completion (SSC) recovers a dense semantic voxel grid from a scan observing 1% of the target volume, under class imbalance beyond 7,000x. We recast SSC as generative semantic scene completion (GSSC): a single discrete-diffusion formulation in three roles. First, paired sparse-dense scene synthesis (PS$^3$) generates matched sparse LiDAR observations with their dense semantic completions, addressing the long tail at its source and yielding the PS$^3$-SemanticKITTI corpus we train on alongside SemanticKITTI. Second, semantic-guided generative scene completion (SGSC) generates the scene from noise with multinomial discrete diffusion, conditioned on the sparse scan through a bird's-eye-view semantic map and a sparse 3D feature stream. Third, the same framework instead refines an existing completion in one flow-matching step: structured source discrete diffusion (S$^2$D$^2$). S$^2$D$^2$ improves the mIoU of SGSC's own output and every external SSC base tested, without base retraining or test-time adaptation. On the strongest base, one step without test-time augmentation reaches 38.8% mIoU on the SemanticKITTI hidden test. To our knowledge that is the best causal, single-sweep, single-sample result on that leaderboard, +2.1 pp over the previous best published score under the same restriction. Four correction steps with eight-view test-time augmentation reach 39.2%, outside that restriction.

Comments18 pages, 12 figures, 4 tables. Supplementary material (29 pages) is included as an ancillary file. Project page: https://shichen.world/GSSC-project-page/ - Code, models and the PS$^3$ dataset: https://github.com/BillyChern/GSSC-S2D2

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑