arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

工业质量控制中的自动焊缝分割:RGB成像与偏振成像及CNN和Transformer架构的对比

Automatic weld seam segmentation for industrial quality control: a comparison of RGB and polarimetric imaging with CNN and transformer architectures

Simone Garbin, Leonardo Venturoso, Marco Todescato

arXiv 2608.25465首次发表:更新:

发表机构

Fraunhofer Italia Research(弗劳恩霍夫意大利研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究对比RGB与偏振成像、CNN与Transformer架构在工业焊缝自动分割中的表现,发现偏振成像无需受控采集即可达高准确率,Transformer在视点偏移下鲁棒性优于CNN。

AI 中文摘要

焊接组件的目视检查仍是许多工业生产流程中自动化程度最低的环节之一,目前仍在很大程度上依赖人工操作者的经验,因此存在操作者间的差异;本研究的研究对象——专用机械舱的制造就是一个典型案例。本研究评估了从RGB成像和偏振成像中进行自动焊缝分割的可行性,对比了受控实验室采集的图像与在真实、不受控条件下采集的图像。在统一的、不依赖阈值的协议下对卷积神经网络(CNN)架构和基于Transformer的架构进行基准测试,为将真实效果与种子噪声区分开,每个CNN使用三个随机种子进行训练。在受控RGB条件下,CNN模型的平均掩码mAP50可达0.87,但在不受控采集条件下降至0.22-0.48,这表明采集设置是检测系统的首要组成部分。采用保留对齐的几何增强的偏振成像可定位先前未见过的焊缝,其平均掩码mAP50可达0.93:与最佳受控RGB结果相当,而非优于,且无需采集控制即可在不受控条件下达到该精度。最明确的架构发现与视点鲁棒性有关:在分布内场景中,Transformer和CNN大致相当;但在测试时视点偏移下,Transformer模型,尤其是RF-DETR,保持高准确率,而所有CNN均失效。该差距在三个种子和分辨率匹配的对照组中均存在,指向架构而非训练分辨率。在CNN家族内,一旦考虑种子方差,模型容量不会带来可靠的分布内增益:小型CNN适用于固定视点,Transformer适用于可变视点。

英文摘要

Visual inspection of welded assemblies remains one of the least automated stages in many industrial production processes, still depending largely on the experience of human operators and thus subject to inter-operator variability; the manufacturing of special-purpose machinery cabins, the setting of this study, is one representative case. This work evaluates the feasibility of automatic weld seam segmentation from RGB and polarimetric imagery, comparing controlled laboratory acquisitions with images captured under real, uncontrolled conditions. Convolutional neural network (CNN) architectures and transformer-based architectures are benchmarked under a unified, threshold-independent protocol, training each CNN with three random seeds to separate genuine effects from seed noise. In controlled RGB conditions, CNN models reach a mean mask mAP50 of up to 0.87, but drop to 0.22-0.48 under uncontrolled acquisition, showing that the acquisition setup is a first-order component of the inspection system. Polarimetric imaging with alignment-preserving geometric augmentation localizes previously unseen welds with a mean mask mAP50 up to 0.93: on par with, rather than ahead of, the best controlled-RGB result, but reaching that accuracy on uncontrolled RGB without requiring acquisition control. The clearest architectural finding concerns viewpoint robustness. In-distribution, transformers and CNNs are broadly comparable; but under a test-time viewpoint shift, the transformer models, and RF-DETR in particular, retain high accuracy while every CNN collapses. The gap holds across three seeds and a resolution-matched control, pointing to architecture rather than training resolution. Within the CNN family, capacity brings no reliable in-distribution gain once seed variance is accounted for: small CNNs suffice for fixed viewpoints, transformers for variable ones.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑