arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10322cs.CV

无坐标几何:LiDAR扩散模型作为3D特征桥梁

Geometry Without Coordinates: LiDAR Diffusion as a 3D Feature Bridge

Samed Doğan, Nico Leuze, Alfred Schöttl

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出以LiDAR为条件的扩散模型,利用2D基础模型伪标签训练,实现从2D先验到3D表示的迁移,并通过线性探针和特征分析验证了其能学习到结构化3D语义特征。

中文摘要 AI 辅助

将大型2D基础模型的丰富先验知识迁移到稀疏的3D LiDAR数据仍然具有挑战性,因为以相当规模训练原生3D基础模型受到数据和标注稀缺性的限制。我们引入了一种以LiDAR为条件的扩散模型,该模型使用来自现成2D基础模型的伪标签进行训练。该模型支持多种输出模态,包括深度、语义分割和实例预测,可通过文本任务提示进行选择。由于模型以LiDAR为条件,其输出和中间UNet特征都可以投影回输入点云,从而能够分析完全在2D监督下学习的3D表示。我们直接在点云空间中研究这种表示,明确排除原始空间坐标以将特征内容与投影几何分离。线性探针在3D语义类别上恢复了高达约23%的平均交并比(MIoU),而匹配的高斯噪声对照组约为3.5%,表明存在大量非平凡结构。跨模态特定特征流的成对余弦相似性揭示了一种分层组织。早期编码器层在跨模态间保持弱对齐但可单独解码,中间层收敛于共享表示,解码器层则重新特化于任务特定输出。这些发现表明,以LiDAR为条件的扩散模型可以仅从2D监督中诱导出结构化的3D表示,其模态依赖的流形在共享瓶颈附近局部统一。这使扩散模型成为将大规模2D先验迁移到稀疏3D领域的可行机制。

英文摘要

Transferring the rich priors of large 2D foundation models to sparse 3D LiDAR remains challenging, as training native 3D foundation models at comparable scale is limited by data and annotation scarcity. We introduce a LiDAR-conditioned diffusion model trained on pseudo-labels from off-the-shelf 2D foundation models. The model supports multiple output modalities, including depth, semantic segmentation and instance prediction, selectable via a textual task prompt. Because the model is conditioned on LiDAR, both its outputs and its intermediate UNet features can be projected back onto the input point cloud, enabling analysis of a 3D representation learned entirely under 2D supervision. We study this representation directly in point-cloud space, explicitly excluding raw spatial coordinates to isolate feature content from projection geometry. Linear probes recover up to ~23% Mean Intersection over Union (MIoU) on 3D semantic classes, compared to ~3.5% for a matched Gaussian-noise control, indicating substantial non-trivial structure. Pairwise cosine similarity across modality-specific feature streams reveals a layered organization. Early encoder layers remain weakly aligned across modalities while individually decodable, intermediate layers converge toward a shared representation, and decoder layers re-specialize toward task-specific outputs. These findings indicate that LiDAR-conditioned diffusion models can induce structured 3D representations from 2D supervision alone, with a modality-dependent manifold that locally unifies near a shared bottleneck. This positions diffusion as a viable mechanism for transferring large-scale 2D priors into sparse 3D domains.

发表机构

  • Munich University of Applied Sciences(慕尼黑应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

↑