AI 中文总结
针对工业场景中故障诊断标注数据稀缺的问题,提出自监督学习方法SAP,通过欠采样振动数据创建折叠频谱训练Transformer重构展开频谱,在CWRU数据集上验证其结合线性探测的效果优于全监督训练。
AI 中文摘要
深度学习是机械故障诊断的新方法,但需要大量标注数据,而标注数据在工业场景中是稀缺资源。我们提出频谱混叠预文本(Spectral Aliasing Pretext,SAP),一种利用频谱混叠对未标注振动数据上的模型进行预训练的自监督学习方法。我们故意对信号欠采样以创建折叠频谱,随后训练Transformer重构原始展开频谱。该预文本任务迫使模型学习机械故障特有的频域不变性,且无需可能具有破坏性的数据增强。在CWRU数据集上的实验表明,SAP能学习到稳定且高度可区分的表示;在线性探测设置中,SAP仅用一小部分标注数据即可快速达到极高的分类性能,且方差低。相比之下,全微调(含全监督训练)并未带来更稳定或更优的结果。总体而言,这些发现表明,在标注数据有限的故障诊断场景中,结合线性探测的SAP比全监督训练更有效、更可靠。
英文摘要
Deep learning is a new way for machinery fault diagnosis but requires extensive labeled data, a scarce resource in industrial settings. We propose Spectral Aliasing Pretext (SAP), a self-supervised learning method that pretrains models on unlabeled vibration data by exploiting spectral aliasing. We deliberately undersample signals to create folded spectrum, then train a Transformer to reconstruct the original unfolded spectrum. This pretext task forces the model to learn frequency-domain invariants characteristic of mechanical faults, without potentially destructive augmentations. Experiments on the CWRU dataset show that SAP learns stable and highly discriminative representations. In a linear probing setting, SAP quickly achieves very high classification performance with only a small fraction of labeled data and low variance. In contrast, full fine-tuning, including fully supervised training, does not lead to more stable or better results. Overall, these findings suggest that SAP combined with linear probing can be more effective and reliable than fully supervised training for fault diagnosis with limited labeled data.