AI 中文总结
研究通过频谱分析检测神经网络故障,介绍自检测神经网络框架SDNN,利用多种技术监测频谱动态,通过课程学习让轻量级检测器网络识别故障模式,实验表明其性能显著优于基于置信度的基线,确立了内部激活频谱分析对神经网络可靠性研究的价值。
AI 中文摘要
神经网络错误分类在内部激活中表现出特征性频谱不稳定性,在输出层不可见。这种现象被识别并形式化为频谱漂移,即连续层激活之间的频域距离,实证验证表明故障的漂移明显高于正确预测。本文介绍了自检测神经网络(SDNN)框架,利用短时傅里叶变换、小波分解和统计矩监测网络深度的频谱动态。通过课程学习,轻量级检测器网络学会识别故障指示模式。在CIFAR-10上的实验表明,SDNN的AUROC达到79.0 +/- 25.3%,显著优于基于置信度的基线。消融研究揭示了小波分解和统计特征的作用,而STFT的作用尚不清楚。这项工作将内部激活的频谱分析确立为神经网络可靠性的一个有前途的方向。
英文摘要
Neural network misclassifications exhibit characteristic spectral instability in internal activations that is invisible at the output layer. This phenomenon is identified and formalized as Spectral Drift -- the frequency-domain distance between consecutive layer activations -- with empirical validation showing that failures exhibit significantly higher drift than correct predictions (1.9% increase, p<0.001). This spectral signature emerges during internal processing but becomes masked in final outputs, explaining why confidence-based detection methods struggle. This work introduces Self-Detecting Neural Networks (SDNN), a framework that monitors spectral dynamics across network depth using Short-Time Fourier Transform, wavelet decomposition, and statistical moments to capture multi-scale spectral features. A lightweight detector network (5% parameter overhead) learns to identify failure-indicative patterns via curriculum learning on progressively challenging distributions: natural misclassifications, distribution shifts, and adversarial perturbations. Experiments on CIFAR-10 demonstrate that SDNN achieves 79.0 +/- 25.3% AUROC across three seeds, substantially outperforming confidence-based baselines including MaxSoftmax (50.5%) and Energy Score (52.9%) by approximately 25-30 percentage points. Ablation studies reveal that wavelet decomposition and statistical features make consistent contributions, while STFT's role remains unclear. This work establishes spectral analysis of internal activations as a promising direction for neural network reliability, revealing diagnostic information inaccessible to output-based approaches.
CommentsSubmitted for ACML 2026 , Under Review