发表机构
University of Asia Pacific(亚太大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对半监督建筑足迹提取中前景背景不平衡导致伪标签偏差的问题,提出双层级类别再平衡框架RBMatch,联合调节伪标签生成与无监督优化,在多个数据集上取得最优性能。
AI 中文摘要
从高分辨率遥感影像中准确提取建筑足迹对城市规划、灾害响应和环境监测至关重要。然而,获取密集的像素级标注成本高昂,这促使人们利用半监督学习(SSL)来利用未标注影像。在遥感中,严重的前景-背景不平衡对自训练构成了特殊挑战,因为它可能使伪标签生成及由此产生的无监督优化偏向多数背景类别。我们表明,仅在一个阶段解决这种不平衡是不够的:仅平衡伪标签选择并不能阻止背景偏差在无监督损失优化期间重新出现,我们将这种失败模式称为“不平衡泄漏”。为解决此问题,我们提出了RBMatch,一种双层级类别再平衡框架,它联合调节伪标签生成和无监督优化。RBMatch将监督学习路径与包含三个组件的自训练模块相结合:用于平衡伪标签选择的自适应类别特定阈值(ACT)、用于减轻无监督损失中类别偏差的置信度感知类别平衡重加权(CACBR),以及用于将预测的未标注数据分布与标注数据先验匹配的分布对齐(DAL)。在WHU、INRIA和马萨诸塞州建筑足迹数据集上,标注比例为1%至10%的实验表明,RBMatch在评估方法中始终取得最佳的建筑IoU和F1分数。在高度不平衡的马萨诸塞州数据集上改进最为显著,在1%标注比例下,RBMatch的IoU比最强基线提高了1.37个百分点,并且是唯一在所有十二个数据集-比例设置中均优于全监督基线的方法。
英文摘要
Accurate building footprint extraction from high-resolution remote sensing imagery is essential for urban planning, disaster response, and environmental monitoring. However, obtaining dense pixel-level annotations is costly, motivating the use of semi-supervised learning (SSL) to leverage unlabeled imagery. In remote sensing, severe foreground--background imbalance poses a particular challenge for self-training, as it can bias pseudo-label generation and the resulting unsupervised optimization toward the majority background class. We show that addressing this imbalance at only one stage is insufficient: balancing pseudo-label selection alone does not prevent background bias from re-emerging during unsupervised loss optimization, a failure mode we term \emph{imbalance leak}. To address this issue, we propose \textbf{RBMatch}, a dual-level class-rebalancing framework that jointly regulates pseudo-label generation and unsupervised optimization. RBMatch combines a supervised learning pathway with a self-training module comprising three components: adaptive class-specific thresholding (ACT) for balanced pseudo-label selection, confidence-aware class-balanced reweighting (CACBR) for mitigating class bias in the unsupervised loss, and distribution alignment (DAL) for matching the predicted unlabeled-data distribution to the labeled-data prior. Experiments on the WHU, INRIA, and Massachusetts building footprint datasets across labeled ratios of 1%--10% show that RBMatch consistently achieves the best building IoU and F1-score among the evaluated methods. The improvement is most pronounced on the highly imbalanced Massachusetts dataset, where RBMatch improves IoU by 1.37 points over the strongest baseline at a 1% labeling ratio and is the only method to outperform the fully supervised baseline across all twelve dataset--ratio settings.