发表机构
Actual Reality Technologies(Actual Reality Technologies)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出利用地面真值缺陷掩码作为空间监督信号,通过注意力对齐损失引导CNN特征图聚焦缺陷区域,结合DDPM增强,在MVTec-AD上显著提升定位性能,且不影响分类精度。
AI 中文摘要
工业检测数据集中的地面真值缺陷掩码通常保留用于评估。本文将其重新用作分类网络训练期间的空间监督信号,教导模型不仅预测什么,而且关注哪里。该方法添加了一个基于激活的注意力对齐损失,将卷积特征图引导至缺陷区域,采用混合监督公式,该公式也适用于没有掩码的样本,例如扩散生成的图像。与DDPM增强相结合,合成图像贡献数量,而掩码贡献空间精度。我们在MVTec-AD瓶子基准上评估了85个模型(四个CNN骨干网络,在五个种子下进行2x2数据/训练因子设计,加上Swin-V2-T变压器基线),定位在排除分类器梯度更新的保留缺陷图像上测量。主要发现:(1)注意力引导训练改善了基于激活的定位(Pixel-AUROC),对于EfficientNetB0(增强)提高了+18.0%(p=0.005,Cohen's d=2.6),对于ResNet50提高了+18.7%(p=0.008),在八个CNN设置中有四个显著(未校正多重比较),分类性能无显著变化;(2)对于EfficientNetB0,数据x训练模式交互显著(p=0.002),与超加性效应一致(组合+13.6% vs 单独效应总和+1.6%);(3)空间表示较弱的架构受益最多,而ConvNeXt-T没有效果,显然是因为其深度卷积激活产生空间信息不足的通道均值图;(4)无监督PatchCore仍然是最强的定位器(Pixel-AUROC=0.983),为监督增益提供了背景。这些结果表明,现有的评估掩码可以作为实用的训练信号,可测量且可重复地改善缺陷分类器的关注位置。
英文摘要
Ground-truth defect masks in industrial inspection datasets are typically reserved for evaluation. This paper repurposes them as spatial supervision signals during training of classification networks, teaching a model not just what to predict but where to look. The method adds an activation-based attention alignment loss that steers convolutional feature maps toward defect regions, in a mixed-supervision formulation that also accommodates samples without masks, such as diffusion-generated images. Combined with DDPM augmentation, synthetic images contribute quantity while masks contribute spatial precision. We evaluate 85 models (four CNN backbones under a 2x2 data/training factorial over five seeds, plus a Swin-V2-T transformer baseline) on the MVTec-AD bottle benchmark, with localization measured on held-out defect images excluded from classifier gradient updates. Main findings: (1) attention-guided training improves activation-based localization (Pixel-AUROC) by +18.0% for EfficientNetB0 with augmentation (p=0.005, Cohen's d=2.6) and +18.7% for ResNet50 (p=0.008), significant in four of eight CNN settings (uncorrected for multiple comparisons) with no significant change in classification; (2) for EfficientNetB0 a data x training-mode interaction is significant (p=0.002), consistent with a super-additive effect (+13.6% combined vs +1.6% summed individual effects); (3) architectures with weaker spatial representations benefit most, whereas ConvNeXt-T shows no effect, apparently because its depthwise-convolution activations yield spatially uninformative channel-mean maps; (4) unsupervised PatchCore remains the strongest localizer (Pixel-AUROC=0.983), contextualizing the supervised gains. These results show that existing evaluation masks can act as practical training signals that measurably and reproducibly improve where defect classifiers attend.
Comments20 pages, 5 figures, 5 tables. Code: https://github.com/Actual-Reality/Glass-Defect-Detection-Attention-Supervision