发表机构
TU München; BMW AG(慕尼黑工业大学; 宝马集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对潜在激光雷达生成中卷积VAE导致的飞行像素问题,提出CRISP扩散解码器,替换后在多数据集上显著降低相关误差,缩小模拟到真实的差距。
AI 中文摘要
潜在激光雷达(LiDAR)流水线存在飞行像素问题:卷积变分自编码器(convolutional VAEs)会模糊尖锐的径向深度不连续性,产生的边缘深度反向投影为悬浮于表面之间的点。我们将此确定为一个主要的、可直接修正的解码器瓶颈,并提出CRISP:一种像素空间扩散解码器,具备与主干网络无关的潜在适配器、基于DiT的去噪器以及支持掩码预测器。CRISP可替代视频VAE和LiDAR专用解码器,同时保持编码器和潜在生成器固定。在KITTI-360、SemanticKITTI和nuScenes数据集上,仅替换解码器可使冻结主干网络的FSVD/FPVD平均降低50.5%;对于通用视频VAE,降幅达71%/74%。在LiDAR专用的LiDM主干网络上,FRID降低71%,且在深度不连续性处增益最大。在预训练的LiDM世界模型中,相同的零样本替换可使FSVD提升15.5%,缩小了模拟到真实的差距。
英文摘要
Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces video-VAE and LiDAR-native decoders alike while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. In a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap.
CommentsAccepted at NeurIPS 2026 (poster). 41 pages, 10 figures, 18 tables. Project page: https://andrea25512.github.io/CRISP/