arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应区域划分与空间监督对比学习下的城市建筑实例分割与细粒度分类

Instance Segmentation and Fine-grained Classification for Urban Buildings with Adaptive Region Dividing and Spatially-Supervised Contrastive Learning

Weiyuan Zhang, Qi Zhang, Hui Huang

arXiv 2609.19631首次发表:更新:

发表机构

Shenzhen University(深圳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对城市点云建筑理解中预定义块划分的局限,提出自适应区域划分与空间监督对比学习的实例分割及细粒度分类方法,在UrbanBIS和STPLS3D上优于现有方法。

AI 中文摘要

在大规模点云中实现对城市建筑的准确实例级和功能理解,对于数字城市建模和城市分析至关重要。然而,城市场景的广泛空间覆盖导致大多数现有方法依赖预定义块进行训练和评估,尽管这种划分在实际应用中很少可用,并且引入了额外的预处理,同时破坏了完整的建筑结构。为解决这一问题,我们提出了一种自适应区域划分策略,并采用统一的场景级评估。具体而言,将3D点云投影到鸟瞰图(BEV)平面上,使用预训练的分割模型检测建筑区域。随后,将检测到的边界框反投影回原始点云,以构建结构对齐的自适应训练块,从而实现无需人工设计的语义引导动态划分。此外,除了实例级理解,很少有方法探索城市建筑的细粒度分类,因此我们还提出了一种带有空间监督对比损失的城市建筑细粒度分类模型。首先,对于每个分割出的建筑,点变换器分类器利用几何、颜色和核心上下文信息联合编码其主体和局部上下文。然后,使用类平衡加权交叉熵来缓解严重的类别不平衡问题。所提出的空间监督对比损失通过为空间邻近、同类别建筑分配更大权重,进一步增强类间可区分性,鼓励紧凑的功能表示,同时分离容易混淆的类别。在UrbanBIS和STPLS3D上的大量实验表明,与现有SOTA方法相比,所提方法在建筑实例分割和细粒度分类方面具有优势。

英文摘要

Accurate instance-level and functional understanding of urban buildings in large-scale point clouds is essential for digital city modeling and urban analysis. However, the extensive spatial coverage of urban scenes leads most existing methods to rely on predefined blocks for training and evaluation, although such partitions are rarely available in real-world applications and introduce additional preprocessing while fragmenting complete building structures. To address this issue, we propose an adaptive region-dividing strategy with unified scene-level evaluation. Specifically, the 3D point cloud is projected onto a bird's-eye-view (BEV) plane, where a pretrained segmentation model is used to detect building regions. The detected bounding boxes are then back-projected to the original point cloud to construct structure-aligned adaptive training blocks, enabling semantically guided dynamic partitioning without manual design. Furthermore, beyond instance-level understanding, few methods have explored fine-grained classification for urban buildings, and thus we also put forward a fine-grained classification model for urban buildings with a spatially-supervised contrastive loss. First, for each segmented building, a point transformer classifier jointly encodes its body and local context using geometric, color, and core-context information. Then, the class-balanced weighted cross-entropy is used to alleviate severe class imbalance. The proposed spatially-supervised contrastive loss further enhances inter-class discriminability by assigning greater weight to spatially proximate, same-category buildings, encouraging compact functional representations while separating easily confused categories. Extensive experiments on UrbanBIS and STPLS3D demonstrate the advantages of the proposed method in building instance segmentation and fine-grained classification compared to existing SOTA methods.

Comments10 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑