溯源而非行为:Edge-IIoTset中的序列化伪影与精准农业入侵检测的无泄漏基准
Provenance, Not Behaviour: A Serialisation Artifact in Edge-IIoTset and a Leakage-Free Benchmark for Precision-Agriculture Intrusion Detection
- Alexandria University(亚历山大大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文发现工业物联网入侵检测基准Edge-IIoTset存在序列化伪影导致标签泄漏,进而构建无泄漏的精准农业入侵检测基准AgriEdge,明确了相关模型的泛化边界与性能表现。
AI中文摘要:
Edge-IIoTset是工业物联网中用于机器学习入侵检测的参考基准,针对该基准的报告结果多集中在99%以上,但本文表明,大部分性能并非来自入侵检测。该数据集附带的预处理流程要求研究人员对7个分类列进行独热编码,其中4个列通过为缺失协议字段编写的占位符拼写,自身就能以1.0000的准确率区分攻击流量与正常流量:数据集构建时,正常流量分支中的字符串“0”与攻击分支中的“0.0”存在差异。标签可从编码文件溯源的序列化伪影中恢复,无需对网络行为进行建模,且能区分两个整理子集的每一行。在5折×3重复交叉验证下,6种标准分类器中有5种达到1.0000±0.0000的准确率,第6种达到0.99998。在修正后的协议下,朴素贝叶斯的宏F1下降0.3005,最强模型的宏F1稳定在0.9503±0.0011。标签编码、序数编码与频率编码的泄漏情况相同。由于整理子集还缺少Modbus和每设备标识,我们基于原始捕获数据通过统一解析重建了基准,生成AgriEdge:共1276122行数据,包含5个具有完整归属的设备,且没有任何列能以高于0.0288的准确率区分类别。留一设备扫描将泛化边界定在感知/执行层,其中随机森林的平衡准确率从0.9988降至0.5083。非独立同分布联邦划分的宏F1损失最多为0.0037,但20轮LoRaWAN训练运行需消耗4.6小时的上行时间。
英文摘要:
Edge-IIoTset is the reference benchmark for machine-learning intrusion detection in the industrial Internet of Things, and results reported on it cluster above 99%. We show that much of that performance is not intrusion detection. The preprocessing recipe distributed with the dataset instructs researchers to one-hot encode seven categorical columns. Four of them separate attack from normal traffic with an accuracy of 1.0000 on their own, through the spelling of the placeholder written for an absent protocol field: the string "0" in the normal-traffic branch of the dataset build against "0.0" in the attack branch. The label is recoverable from a serialisation artifact encoding file provenance, with no network behaviour modelled, and separates every row of both curated subsets. Under 5-fold x 3-repeat cross-validation, five of six standard classifiers attain exactly 1.0000 +/- 0.0000 accuracy and the sixth attains 0.99998. Under a corrected protocol, naive Bayes falls by 0.3005 macro-F1 and the strongest model settles at 0.9503 +/- 0.0011. Label, ordinal and frequency encoding leak identically. Because the curated subsets also lack Modbus and per-device identity, we rebuild the benchmark from the raw captures under uniform parsing, producing AgriEdge: 1,276,122 rows, five devices with full attribution, and no column separating the classes above 0.0288. A leave-one-device-out sweep locates the generalisation boundary at the perception/actuation layer, where random forest falls from 0.9988 to 0.5083 balanced accuracy. Non-IID federated partitioning costs at most 0.0037 macro-F1, but a 20-round LoRaWAN training run costs 4.6 hours of uplink.