工业控制系统(ICS)网络安全数据集:关于覆盖范围、评估实践与结构缺口的系统元综述
ICS Cybersecurity Datasets: A Systematic Meta-Review of Coverage, Evaluation Practice, and Structural Gaps
浏览论文内容
中文总结 AI 辅助
本文通过元综述分析83个ICS网络安全数据集的结构偏差与评估实践缺陷,确定三种结构失衡并提出含跨阶段数据集构建等内容的协调研究议程。
中文摘要 AI 辅助
工业控制系统(ICS)的入侵检测研究高度依赖公开数据集,但尚无研究系统评估现有数据集集合是否支持当前的评估主张。本文针对2019年至2026年间的18项研究开展元综述,从中识别出83个ICS或与ICS直接相关的网络安全数据集,采用统一的五维分类法对这些数据集进行协调与特征刻画。该分类法揭示数据集集合存在结构偏差:85.5%的数据集集中于后期运营技术(OT)破坏策略,跨阶段信息技术(IT)/运营技术(OT)演进序列仅占8.4%,普渡层级(Purdue hierarchy)0级的现场设备证据基本缺失,运营来源的数据仅占集合的15.7%。对评估实践的并行审计显示,没有研究报告采用流式评估,不到一半的研究采用规范的训练/测试划分,仅有两项研究满足可复现性要求。此外,分类法-评估耦合分析表明,数据集不平衡会限制多项评估实践的范围与可行性。基于这些发现,本文确定了三种结构失衡:i)架构浅层化,ii)演进压缩,iii)跨域替代,并提出协调研究议程,涵盖跨阶段数据集集合构建、时间结构化基准测试、事件级标签标准以及运营数据共享的治理框架。
英文摘要
Intrusion detection research in Industrial Control Systems (ICS) heavily depends on public datasets, yet no prior work has systematically assessed whether the collective dataset corpus supports current evaluation claims. This paper addresses this gap through a meta-review of 18 studies between 2019 and 2026, from which 83 ICS, or ICS directly related, cybersecurity datasets are identified, harmonised, and characterised using a unified five-dimensional taxonomy. The taxonomy reveals that the corpus is structurally skewed: 85.5% of datasets concentrate on late-stage OT Disruption tactics, cross-stage IT/OT progression sequences are present in only 8.4% of cases, field-device evidence at Level 0 of the Purdue hierarchy is effectively absent, and operationally sourced data accounts for only 15.7% of the collection. A parallel audit of evaluation practices shows that zero report streaming evaluation, fewer than half apply disciplined train/test partitioning, and only two satisfy reproducibility requirements. Furthermore, a taxonomy-evaluation coupling analysis shows that dataset imbalances constrain the scope and feasibility of several evaluation practices. Based on these findings, we identify three structural imbalances: i) architectural shallowness, ii) progression compression, and iii) cross-domain substitution, and derive a coordinated research agenda which covers cross-stage corpus construction, temporally structured benchmarking, event-level label standards, and governance frameworks for operational data sharing.