arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23547cs.CRcs.LG

训练时数据污染下工业控制系统异常检测模型的鲁棒性

Robustness of Anomaly Detection Models for Industrial Control Systems under Training-Time Data Contamination

  • Ontario Tech University(安大略理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Mustafa Umut Ozbek, Taiwo Ojo, Pooria Madani, Khalil El-Khatib, Li Yang

AI总结:

本文在 SWaT 基准上评估训练时数据污染下 11 种异构 ICS 异常检测器的鲁棒性,发现鲁棒性具强模型依赖性,注入型污染危害更大,PCA 等模型较稳定,凸显训练数据完整性的重要性。

AI中文摘要:

基于机器学习的异常检测正越来越多地应用于工业控制系统(ICS),但多数研究假设检测器的训练数据是可信的。实际场景中,训练数据可能因日志遭入侵、标注错误、 historian 记录被篡改或不安全的重训练流程而被污染。本文在安全水处理(SWaT)基准上评估离线 ICS 异常检测流水线在训练时污染下的鲁棒性,针对三种污染策略(随机注入、相似度目标注入、特征噪声注入)评估 11 种异构异常检测器:前两种向标称训练池中插入攻击样本,第三种向选定的正常训练样本添加有界高斯噪声;这些攻击是基于污染而非梯度驱动的投毒方法。在统一离线协议下,使用干净的验证集和测试集评估 1%至 10%的污染预算。结果表明,鲁棒性具有强模型依赖性,无法仅通过干净数据的性能预测;基于注入的污染造成的性能下降最大,尤其对局部密度和距离型检测器,而特征噪声污染的影响相对有限;PCA、SVM、HBOS 和 IForest 保持相对稳定,调优后的神经检测器表现出中等鲁棒性。总体而言,研究结果强调了在基于机器学习的 ICS 监控中训练数据完整性的重要性,该结论适用于所评估的数据集、模型和威胁假设。

英文摘要:

Machine-learning-based anomaly detection is increasingly used in industrial control systems (ICS), yet most studies assume that detector training data is trustworthy. In practice, training data may be corrupted through compromised logs, labeling errors, manipulated historian records, or unsafe retraining processes. This paper evaluates the robustness of offline ICS anomaly-detection pipelines on the Secure Water Treatment (SWaT) benchmark under training-time contamination. We assess 11 heterogeneous anomaly detectors under three contamination strategies: random injection, similarity-targeted injection, and feature-noise injection. The first two insert attack samples into the nominal training pool, while the third adds bounded Gaussian noise to selected normal training samples. These attacks are contamination-based rather than gradient-driven poisoning methods. Contamination budgets from 1% to 10% are evaluated using clean validation and test sets under a unified offline protocol. The results show that robustness is strongly model-dependent and cannot be predicted from clean-data performance alone. Injection-based contamination causes the greatest degradation, particularly for local-density and distance-based detectors, whereas feature-noise contamination has a comparatively limited effect. PCA, SVM, HBOS, and IForest remain relatively stable, while the tuned neural detectors demonstrate intermediate robustness. Overall, the findings highlight the importance of training-data integrity in ML-enabled ICS monitoring, subject to the evaluated dataset, models, and threat assumptions.

补充信息

↑