arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CW-BASS v2:基于基础模型教师的半监督分割中感知饱和度的伪标签选择

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

Ebenezer Tarubinga

arXiv 2608.12773首次发表:更新:

发表机构

Ebenworks Systems(埃本沃克斯系统公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CW-BASS v2是一种感知饱和度的伪标签选择方法,它针对DINOv2教师模型的置信饱和问题,结合预留校准等技术,在6个基准数据集上恢复UniMatch V2操作点并提升性能。

AI 中文摘要

长期以来,半监督语义分割始终围绕一个核心问题:应信任哪些伪标签?过往一系列选择规则(如动态阈值、按类课程学习、软置信度权重等)已针对当时噪声大、置信度低的ResNet教师模型解决了该问题。而自监督基础编码器改变了这一局面:使用DINOv2教师模型时,置信度会出现饱和现象,因此适用于弱教师的过滤方法可能会损害强教师的性能。我们提出了CW-BASS v2,这是一种感知饱和度的伪标签选择方法,它能读取教师模型的置信度状态,而非固守单一规则。该方法结合了预留校准、无偏按类噪声估计,以及可证明将保留率限制在1以下的自适应置信度下限,并将它们整合为一个单步门控机制:在预留子集上测量教师置信集的可靠性,即pi_kept = Pr[正确 | c >= tau],当满足所需置信度(pi_kept >= tau)时严格过滤,否则退回到自适应下限。该边界是预先存在的操作阈值,而非针对mIoU调整的值,在6个DINOv2教师模型上,它能盲选正确的严格过滤或下限策略。CW-BASS v2因此在饱和基准上恢复了UniMatch V2的操作点,例如在Pascal VOC 1/8数据集上,其结果为87.4,接近报告的87.9;在Cityscapes数据集上误差在0.5以内;在置信集不可靠的地方(pi_kept约为89%,ADE20K数据集)表现更优,在自适应下限占优的地方(单种子情况下mIoU提升1.5)也有改善。该门控机制是有原则的,因为它避免的失败是可测量而非假设的:在可靠的饱和教师模型上,置信度分布的动态范围会崩溃(Pascal数据集98%的像素置信度>=0.95),因此自适应截止会使保留掩码泛滥,自训练会退化为确认偏差。

英文摘要

Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident ResNet teachers of their day. Self-supervised foundation encoders change the regime: with a DINOv2 teacher, confidence saturates, so the filtering that helped a weak teacher can hurt a strong one. We propose CW-BASS v2, a saturation-aware pseudo-label selection method that reads the teacher's confidence regime rather than committing to one rule. It pairs held-out calibration, an unbiased per-class noise estimate, with a self-adaptive confidence floor that provably bounds retention away from 1, and combines them in a one-pass gate: measure the reliability of the teacher's confident set, pi_kept = Pr[correct | c >= tau], on a held-out slice, and filter strictly when it meets the confidence demanded (pi_kept >= tau), falling back to the adaptive floor otherwise. The boundary is the pre-existing operating threshold, not a value tuned to mIoU, and across six DINOv2 teachers it makes the correct strict-vs-floor call blind. CW-BASS v2 thus recovers the UniMatch V2 operating point on the saturated benchmarks by selecting strict (Pascal VOC 1/8 87.4 against its reported 87.9; Cityscapes within 0.5), and improves on it where the confident set is unreliable (pi_kept ~ 89%, ADE20K), where the floor edges ahead (+1.5 mIoU, single seed). The gate is principled because the failure it avoids is measured, not assumed: on a reliable, saturated teacher the confidence distribution's dynamic range collapses (98% of Pascal pixels >= 0.95), so an adaptive cutoff floods the retention mask and self-training decays into confirmation bias.

CommentsSubmitted to IEEE TPAMI. 22 pages, 11 figures, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑