arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Polis:城市规模的3D自监督学习

Polis: 3D Self-Supervision at City Scale

Alexander Rusnak, Sophia Kovalenko, Jingru Wang, Ismail Moudden, Xiru Wang, Frédéric Kaplan

arXiv 2608.29426首次发表:更新:

发表机构

École Polytechnique Fédérale de Lausanne(洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对城市规模3D场景设计的Polis,结合SIGReg等技术,在14个城市与建筑规模基准上,提升了语义表示性能,同时展现了该领域自监督学习专业化的局限。

AI 中文摘要

从城市规模的3D模型中提取的可靠语义表示对城市分析、基础设施监测、自主系统和遗产保护日益重要。然而,航空测量捕获的大范围城市场景与大多数3D自监督模型预训练所用的室内、对象级和自动驾驶LiDAR数据存在显著差异。我们推出Polis,据我们所知,这是将Sketched Isotropic Gaussian Regularization(SIGReg,草图式各向同性高斯正则化)作为原生点云编码器目标的首次应用,并通过涵盖14个城市和建筑规模语料库的冻结特征基准对其进行评估。Polis结合了几何匹配的余弦不变性、SIGReg以及VICReg风格的反坍缩项,采用包含12800个场景的室外预训练混合数据集和保持重力的空间视图采样。控制消融实验表明,该目标在相同的代表性室外语料库上,优于学生-教师架构替代方案以及不含反坍缩项的Polis版本。在3个与预训练不重叠的城市数据集上,Polis在高容量冻结探测下达到23.8%的平均mIoU,而次优编码器为16.3%;在匹配的点和体素预算下,分别为17.3%和16.1%。在预训练中见过训练集的数据集上,同样保持城市规模的领先优势。在具有细粒度立面和街景标签的局部地面捕获数据上,排名发生反转。我们的结果表明,分布正则化的联合嵌入架构可在具有挑战性的城市规模3D场景中取得成功,且当自监督学习针对该领域的捕获几何和空间上下文设计时,迁移性能会提升,同时也揭示了这种专业化的局限性。

英文摘要

Reliable semantic representations derived from city-scale 3D models are increasingly important for urban analysis, infrastructure monitoring, autonomous systems, and heritage conservation. However, urban scenes of large spatial extent captured through aerial surveying differ substantially from the indoor, object-level, and self-driving LiDAR data used to pretrain most 3D self-supervised models. We introduce Polis, to our knowledge the first application of Sketched Isotropic Gaussian Regularization (SIGReg) as an objective for a native point cloud encoder, and evaluate it through a frozen-feature benchmark spanning fourteen city- and building-scale corpora. Polis combines geometrically matched cosine invariance, SIGReg, and VICReg-style anti-collapse terms with a 12.8k-scene outdoor pretraining mixture and gravity-preserving spatial view sampling. Controlled ablations show that this objective outperforms student--teacher architecture alternatives, as well as Polis versions without anti-collapse terms, on the same representative outdoor corpus. On three pretraining-disjoint city datasets, Polis reaches $23.8\%$ mean mIoU versus $16.3\%$ for the next-best encoder under high-capacity frozen probing, and $17.3\%$ versus $16.1\%$ at a matched point and voxel budget. The same city-scale lead holds on datasets whose training sets were seen in pretraining. On localized terrestrial captures with fine-grained facade and streetscape labels, the ranking reverses. Our results show that distributionally-regularized joint embedding architectures can be successful on challenging city-scale 3D scenes, and that transfer improves when self-supervision is designed for the capture geometry and spatial context of this domain while also revealing the limits of this specialization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑