发表机构
Cheriton School of Computer Science; University of Waterloo(切里顿计算机科学学院; 滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对弱监督语义分割中边界精度不足问题,提出DS-CRF框架,将软伪标签与SAM边界解耦为不同监督信号,在MS COCO上达到56.5% mIoU的新纪录。
AI 中文摘要
弱监督语义分割(WSSS)从图像级标签学习像素级预测。近期工作侧重于改进从大型视觉-语言模型(通常为CLIP)中提取的粗糙CAM,但对其在分割边界处的精度提升甚微。该任务被委托给诸如DenseCRF的后处理方法。然而,由于DenseCRF依赖低级颜色线索,当相邻像素颜色相似时,它可能将正确标签翻转为错误标签。SAM最近被作为自然替代方案采用,但在先前工作中,它仅承担DenseCRF的角色,作为输出独热伪标签的中间“细化”步骤。通过丢弃CAM中有价值的不确定性,这些独热伪标签将边界错误转化为自信的错误目标。我们的关键见解是,CAM应与其各自的损失一起监督训练,同时利用SAM边界,而不是将它们融合为单一硬目标。受CRF势能启发,我们提出一个框架,将软伪标签解耦为一元监督,将二值边缘图解耦为成对监督。我们通过单阶段模型DS-CRF实现该框架,使用来自此http URL的CAM和来自SAM的边界。DS-CRF在MS COCO上达到了56.5% mIoU的新最先进水平。
英文摘要
Weakly Supervised Semantic Segmentation (WSSS) learns pixel-level predictions from image-level tags. Recent work focuses on improving coarse CAMs extracted from large vision-language models (commonly CLIP), but does little to improve their accuracy along segment boundaries. That job is instead delegated to a post-processing method like DenseCRF. However, because DenseCRF relies on low-level colour cues, it can flip correct labels to incorrect ones when neighbouring pixels share similar colours. SAM has recently been adopted as a natural alternative, yet it simply takes on DenseCRF's role as an intermediate "refinement" step that outputs one-hot pseudo-labels in prior work. By discarding the valuable uncertainty in CAMs, these one-hot pseudo-labels turn borderline errors into confidently wrong targets. Our key insight is that CAMs should supervise training alongside SAM boundaries, each through its own loss, rather than being fused together into a single hard target. Inspired by CRF potentials, we propose a framework that disentangles soft pseudo-labels as unary supervision and binary edge maps as pairwise supervision. We realize our framework in a single-stage model, DS-CRF, using CAMs from dino.txt and boundaries from SAM. DS-CRF sets a new state-of-the-art of 56.5% mIoU on MS COCO.