arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02821cs.CR

检测器能看到什么:独立于决策规则评估网络物理系统(CPS)异常检测器

Where Does Detection Fail? A Residual-Evidence Diagnostic Workflow for Industrial Control Systems

Peiran Shi, Jian Xiang, Xiang Zhang, Chenglong Fu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出不依赖决策规则的CPS异常检测器评估方法,用归一化残差能量评估阶段1,在三个基准上验证了不同检测器性能差异,可区分检测失败的不同原因。

中文摘要 AI 辅助

异常检测器通常是网络物理系统(CPS)的最后一道防线。但从深度神经网络到不变量模板等不同方式构建的检测器,通常使用单一工作点下的精确率、召回率或F1值进行比较。这些分数混合了两个独立的方面:检测器对物理过程的表征效果,以及其警报阈值的设置效果。因此,我们将CPS异常检测器视为一个两阶段流水线:阶段1将观测值映射到残差,阶段2将残差映射到警报。我们不仅对最终警报评分,而是使用归一化残差能量直接评估阶段1,该指标与训练正态参考分布的Kullback-Leibler散度存在精确关联。由于它不依赖特定警报规则,可单独测量攻击分离度、跨训练-测试差距的稳定性,以及检测器对被控对象编码的紧凑性。无需对每个检测器进行调优,我们将此评估应用于五个检测器——GDN、FuSAGNet、TranAD、NSIBF和GeCo——涉及三个CPS基准:SWaT、WADI和HAI。尽管这些检测器在SWaT上的ROC-AUC值相近,但在相同误报率下性能差异超过一个数量级。排名也会随测试平台变化:TranAD在HAI上排名第一但在SWaT上排名最后,而NSIBF在WADI上排名第一但在HAI上排名最后。在WADI上,局部攻击可规避汇集所有通道证据的检测器,这有助于解释为何NSIBF优于在其他基准上表现良好的方法。这些结果表明,检测失败可能源于不同原因:表征薄弱、阈值校准不佳,或攻击几乎无物理影响。无决策规则的分析有助于区分这些原因。

英文摘要

Residual-based anomaly detectors for industrial control systems (ICS) convert model residuals into anomaly scores and alarms, but final detection metrics provide little guidance on which component to examine when performance is poor. We present a retrospective diagnostic workflow built around covariance-normalized residual energy (CNRE), a measurement based on squared Mahalanobis distance. The workflow evaluates whether the normal reference still describes the evaluation recording, whether attacks separate from normal operation under a specified reference, and how scoring and alarm rules use fixed residuals. Across fifteen detector--dataset pairs spanning five detectors on SWaT, WADI, and HAI, reference updates improve attack--normal separation in some cases but degrade it in others. A retrospective policy that updates the reference only when recommended by a label-free screening rule increases mean AUC from 0.751 with no update and 0.759 with unconditional updating to 0.809. On the twelve non-GDN pairs, which did not inform the screening boundaries, the corresponding AUC is 0.805, compared with 0.759 and 0.747. Rescoring CNRE under a fixed reference reproduces most of the AUC gain from a reference update on WADI, but not on SWaT. A detailed GDN case study further connects the diagnostic stages. Normal-only adaptation reduces normal prediction error by 27.3\%, yet its effect on attack separation depends on the reference protocol. With residuals fixed, improved score ranking also does not guarantee that normal-calibrated alarms preserve their nominal false-positive rate. Together, these results support diagnosing reference fit, attack evidence, and alarm calibration separately when selecting intervention stages and interpreting their effects.

补充信息

↑