发表机构
Munich University of Applied Sciences(慕尼黑应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出利用基础模型引导的扩散方法恢复高线束 LiDAR 分辨率,在极稀疏输入下显著优于插值,并量化了恢复与分辨率的权衡。
AI 中文摘要
高线束 LiDAR 传感器成本高昂,然而许多感知流程需要密集的角采样。我们以预训练的 Stable Diffusion 模型为骨干,利用来自 2D 基础模型的伪深度目标,微调了一个以 LiDAR 为条件的深度模型。训练过程中,LiDAR 条件在不同线束预算下被随机抽取。随后,我们研究了从严重抽取的输入中能恢复多少 LiDAR 扫描,并刻画了在不同输入线束预算下的性能表现。我们在 nuScenes 数据集上针对物理上留出的真实线束进行评估,并将恢复结果与拟合精度分开报告。我们的模型在极稀疏场景下优势最大,从 4 线束输入中实现了 66.8% 的 δ1.25 精度,而散射插值仅达到 45.1%。分层误差分解进一步揭示,平面表面最先恢复,而引入深度不连续性的物体最早退化。综合这些结果,量化了基础模型引导的 LiDAR 增强中恢复与分辨率之间的权衡。
英文摘要
High-beam-count LiDAR sensors are costly, yet many perception pipelines require dense angular sampling. Using a pretrained Stable Diffusion model as the backbone, we fine-tune a LiDAR-conditioned depth model with pseudo-depth targets from a 2D foundation model. During training, the LiDAR conditioning is randomly decimated at different beam budgets. We then investigate how much of a LiDAR scan can be recovered from heavily decimated input and characterize performance across the input beam budget. We evaluate against physically held-out real beams on nuScenes and report recovery separately from fit accuracy. Our model yields its largest advantage in very sparse regimes, achieving a $δ_{1.25}$ accuracy of $66.8$% from $4$-beam input where scattered interpolation reaches only $45.1$%. A class-stratified error breakdown further reveals that planar surfaces recover first while objects introducing depth discontinuities degrade earliest. Together, these results quantify the recovery/resolution trade-off for foundation-model-guided LiDAR enhancement.