发表机构
Institute of Science Tokyo(东京理科大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对基于自动编码器的异常声音检测问题,提出MemNMF方法,该方法基于线性预测编码频谱,通过初始化内存模块并进行注意力加权组合来重建输入,实验表明其能改进自动编码器基线,在复杂条件下鲁棒性强。
AI 中文摘要
基于自动编码器的异常声音检测对机器状态监测很有吸引力,因其仅使用正常录音进行训练,并根据重建误差产生可解释的异常分数。大多数先前工作使用频谱图自动编码器,但重建详细的时频模式对噪声和瞬变敏感,且模型能很好地重建一些异常输入,削弱了正常与异常的分离。我们提出MemNMF,一种在线性预测编码频谱上运行的约束重建方法。MemNMF从在正常LPC频谱上学习的NMF字典初始化内存模块,并将每个输入重建为典型正常频谱模式的注意力加权组合。在多种机器类型和运行条件下对MIMII和DCASE 2020任务2进行的实验表明,LPC频谱输入改进了标准自动编码器基线,且MemNMF进一步提升,在噪声、非平稳设置下具有很强的鲁棒性。
英文摘要
Autoencoder-based anomalous sound detection is attractive for machine condition monitoring because it can be trained using only normal recordings and yields an interpretable anomaly score from reconstruction error. Most prior work uses spectrogram autoencoders, but reconstructing detailed time--frequency patterns is sensitive to noise and transients, and models can reconstruct some anomalous inputs well, weakening normal--anomaly separation. We propose MemNMF, a constrained reconstruction method that operates on the Linear Predictive Coding spectrum, a compact estimate of the spectral envelope. MemNMF initializes a memory module from an NMF dictionary learned on normal LPC spectra and reconstructs each input as an attention-weighted combination of prototypical normal spectral patterns. Experiments on MIMII and DCASE 2020 Task 2 across multiple machine types and operating conditions show that LPC-spectrum inputs improve a standard autoencoder baseline and that MemNMF yields further gains, with especially strong robustness under noisy, non-stationary settings.