发表机构
Faculty of Engineering, University of Deusto(德乌斯托大学工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对自适应网络入侵检测系统升级时的候选可比性问题,在三个公开数据集上开展实验,发现挑战者构建与证据量会影响升级效果,验证可改善证据劣势挑战者的表现,明确需控制相关因素以确保评估可靠性。
AI 中文摘要
自适应网络入侵检测系统会在漂移警报后重新训练分类器,但警报仅检测到变化,并未确定挑战者应替换已部署的 incumbent( incumbent 指当前运行的模型)。升级与安全相关,因为它会改变后续攻击检测的负责模型,且评估升级存在方法论问题:升级结论可能取决于挑战者的构建方式及支持它的证据量。我们使用自包含挑战者流水线、嵌套候选大小控制、对9种更新策略的通用框架比较,以及将每个精确特征向量限制在训练或探测角色的最终敏感性设置,在 CICIDS2017、UNSW-NB15 和 ToN-IoT 数据集上测试了这种依赖性。incumbent 拥有的冻结预处理放大了表观升级危害;使用自包含挑战者流水线时,平均全漂移危害未持续存在。将每个类的名义候选证据从512个样本增加到2000个,在池构建的渐进漂移下,平衡准确率分别提高了+0.53、+1.67和+0.38个点:在三个基准中均为正且统计上可分辨,但在物质上依赖基准而非均匀,且主要由更少的假阳性驱动。策略结论部分稳健:策略排序随候选可比性变化,无全局主导策略,且无标签估计器与校准集成的早期兼容性陈述范围缩小。验证帮助了证据劣势的挑战者,但在均等情况下未增加平均收益。对真实时间顺序流量的13次重放显示,始终部署无净危害。因此,评估升级时应明确控制、报告和解释挑战者构建与证据。
英文摘要
Adaptive network intrusion detection systems retrain classifiers after drift alarms, but an alarm detects change; it does not establish that a challenger should replace the deployed incumbent. Promotion is security-relevant because it changes the model responsible for subsequent attack detection, and evaluating it has a methodological problem: promotion conclusions may depend on how the challenger was constructed and on how much evidence supports it. We test that dependence on CICIDS2017, UNSW-NB15 and ToN-IoT with self-contained challenger pipelines, nested candidate-size controls, a common-harness comparison of nine update policies, and a final sensitivity confining every exact feature vector to one evaluation, training or probe role. Incumbent-owned frozen preprocessing amplified apparent promotion harm; with self-contained challenger pipelines the mean full-drift harm did not persist. Raising nominal candidate evidence from 512 to 2,000 samples per class improved promotion under pool-constructed progressive drift by +0.53, +1.67 and +0.38 balanced-accuracy points: positive and statistically resolved in all three benchmarks, but materially benchmark-dependent rather than homogeneous, and driven mainly by fewer false positives. Policy conclusions were partially robust: policy ordering changed with candidate comparability, no policy globally dominated, and earlier compatibility statements for a label-free estimator and a calibrated ensemble narrowed. Validation helped evidence-disadvantaged challengers but added no average benefit at parity. Thirteen replays on real, time-ordered traffic showed no net harm from always deploying. Challenger construction and evidence should be controlled, reported and interpreted explicitly when promotion is evaluated.
Comments29 pages, 3 figures, 11 tables. Supplementary material (Online Resource 1) is included as an ancillary file