发表机构
Tongji University(同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对工业异常检测现有指标的缺陷,提出结合无分布上限容差阈值与配对减交叉空间检验的方法,在多数据集上验证了其有效性,支持同时报告工作点性能与经机会校正的空间证据。
AI 中文摘要
工业视觉检测研究通常报告受试者工作特征曲线下面积(AUROC)以及异常图与缺陷掩码的重叠度。这两种指标均未指定选定阈值下的虚警率,而重复的缺陷位置和掩码几何形状会夸大重叠度。我们将无分布上限容差阈值与配对减交叉空间检验相结合,该检验比较每个检测器的得分贡献位置与匹配缺陷掩码及其他图像掩码;率的差值定义为相对于经验机会重叠率的空间证据提升。我们在来自三种ISP-AD模态的120张点缺陷图像以及三种固定数据划分上评估了三种检测器。在378个警报中,230个与匹配掩码重叠。然而,在9个检测器-模态单元中的8个里,配对率与交叉率相似;仅DINOv2-ASM具有正的95%自助法下界(提升值0.259,区间0.159-0.347)。在独立的Magnetic Tile Defect数据集上,相同分析得到Wide ResNet-50(WRN50)补丁记忆的提升值为0.203(区间0.169-0.236),Vision Transformer B/16(ViT-B/16)补丁记忆的提升值为0.231(区间0.202-0.262),两者的单侧置换p值均为10^-5。当交叉掩码限定为同一缺陷类别时,提升值仍为0.185和0.210。精确样本规划表明,使用150个校准正常样本时,仅当目标假阳性率为1.98%或更高时,才能支持95%置信度的无分布声明;1%的目标则需要至少299个正常样本。研究结果支持在报告AUROC和原始掩码重叠度的同时,报告工作点性能和经机会校正的空间证据。
英文摘要
Studies of industrial visual inspection commonly report the area under the receiver operating characteristic curve (AUROC) and the overlap between anomaly maps and defect masks. Neither measure specifies the false-alarm rate at a selected threshold, while recurrent defect locations and mask geometry can inflate overlap. We combine a distribution-free upper tolerance threshold with a paired-minus-crossed spatial test. This test compares each detector's score-contributing locations with the matched defect mask and with masks from other images; the difference in rates defines spatial-evidence lift relative to the empirical chance-overlap rate. We evaluate three detectors on 120 point-defect images from three ISP-AD modalities and three fixed data splits. Of 378 alarms, 230 overlap the matched mask. Paired and crossed rates are nevertheless similar in eight of nine detector--modality cells; only DINOv2--ASM has a positive 95\% bootstrap lower bound (lift 0.259, 95\% interval 0.159--0.347). On the independent Magnetic Tile Defect dataset, the same analysis gives lifts of 0.203 (0.169--0.236) for Wide ResNet-50 (WRN50) patch memory and 0.231 (0.202--0.262) for Vision Transformer B/16 (ViT-B/16) patch memory, with one-sided permutation $p=10^{-5}$ for both. When crossed masks are restricted to the same defect class, the lifts remain 0.185 and 0.210. Exact sample planning shows that, with 150 calibration normals, a 95\%-confidence distribution-free claim is supported only for target false-positive rates of 1.98\% or higher; a 1\% target requires at least 299 normals. The results support reporting operating-point performance and chance-corrected spatial evidence alongside AUROC and raw mask overlap.
Comments15 pages, 3 figures, 9 tables