AI 中文总结
针对钢表面缺陷识别中标记数据稀缺问题,采用基于Transformer的掩码自动编码器,通过随机掩码部分图像块让轻量级解码器重建,联合辅助缺陷定位目标训练编码器,聚类预训练编码器特征,取得较高匹配准确率。
AI 中文摘要
钢表面缺陷的自动视觉检测是一项反复出现的质量控制任务,其中带标签的缺陷数据稀缺且获取成本高,而未标记的表面图像丰富,这促使了无需类别标签就能学习有用表示的自监督方法。本文使用基于Transformer的掩码自动编码器来学习钢表面缺陷的表示以进行无监督分组。在预训练期间,75%的输入图像块被随机掩码,一个轻量级解码器从可见的25%重建掩码区域。编码器与辅助缺陷定位目标联合训练,仅用作训练信号而不作为检测器评估。解码器的结构相似性得分达到0.92,均方误差为0.47。然后使用UMAP进行降维和凝聚聚类对预训练编码器的特征进行聚类,针对六种已知缺陷类别,匈牙利匹配准确率达到91.3%。
英文摘要
Automated visual inspection of steel surface defects is a recurring quality control task in which labeled defect data is scarce and costly to obtain, while unlabeled surface images are abundant, which motivates self supervised methods that learn useful representations without class labels. A Transformer based Masked Autoencoder is used here to learn representations of steel surface defects for unsupervised grouping. During pretraining, 75% of the input image patches are randomly masked, and a lightweight decoder reconstructs the masked regions from the visible 25%. The encoder is trained jointly with an auxiliary defect localization objective, used only as a training signal and not evaluated as a detector. The decoder reaches a structural similarity score of 0.92 and a mean squared error of 0.47. Features from the pretrained encoder are then clustered using UMAP for dimensionality reduction and Agglomerative clustering, reaching a Hungarian matched accuracy of 91.3% against the six known defect categories.