多类别、多层级网络入侵检测:一个全面且可复现的基准
Multi-Class, Multi-Tier Network Intrusion Detection: A Comprehensive and Reproducible Benchmark
- RENCI, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校文艺复兴计算研究所)
- Fordham University(福特汉姆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一个全面且可复现的网络入侵检测基准,修正了CIC-IDS2017的标注错误,在三个嵌套级别评估多种模型,软投票集成达到最优细粒度宏F1 0.955,并提供开源可配置流程。
AI中文摘要:
机器学习(ML)和深度学习(DL)近年来主导了入侵检测系统(IDS)的研究。不幸的是,许多现有研究由于在ML和DL流程中的关键疏忽和错误,从数据收集和标注到特征工程以及模型训练和评估,产生了夸大的结果和不可靠的基准。CIC-IDS2017是网络入侵检测的标准基准。然而,由于标注错误、不一致的流提取、潜在的泄漏以及以良性流量为主的性能评估指标,该数据集上的许多已发表结果难以比较。在本文中,我们提出了一个具有修正的PCAP级标注和包含多种ML模型的完整评估流程的全面基准。我们在三个嵌套级别上评估了十一个表格分类器:二元攻击检测、九类攻击家族归属和十五类细粒度分类。随机森林、XGBoost和LightGBM的软投票集成在细粒度层级上获得了最佳的宏F1分数0.955,粗粒度和二元层级的宏F1分数分别为0.980和0.999。我们进一步基于特征重要性分析进行了特征选择研究。这个全面的基准流程是可配置且开源的,支持针对新数据集的新特征提取和模型插件。未来的工作应将此流程作为参考点,用于更丰富的特征、稀有类别分析以及模型向新数据集和攻击类别的泛化。
英文摘要:
Machine learning (ML) and deep learning (DL) have dominated Intrusion Detection System (IDS) research in recent years. Unfortunately, many existing studies have produced inflated results and unreliable benchmarks due to critical oversights and mistakes in the ML and DL pipeline, from data collection and labeling to feature engineering and model training and evaluation. CIC-IDS2017 is a standard benchmark for network intrusion detection. Still, many published results on this dataset are difficult to compare due to labeling errors, inconsistent flow extraction, potential leakage, and performance evaluation metrics dominated by benign traffic. In this paper, we present a comprehensive benchmark with corrected PCAP-level labeling and a complete evaluation pipeline with diverse ML models. We evaluate eleven tabular classifiers at three nested levels: binary attack detection, nine-class attack-family attribution, and fifteen-class fine-grained classification. A soft-voting ensemble of Random Forest, XGBoost, and LightGBM obtains the best fine-tier macro-F1 of 0.955, with coarse and binary macro-F1 scores of 0.980 and 0.999, respectively. We further conducted a feature selection study based on an analysis of feature importance. This comprehensive benchmark pipeline is configurable and open-source, enabling new feature extraction and model plugins for new datasets. Future work should use this pipeline as a reference point for richer features, rare-class analysis, and model generalization towards new datasets and attack classes.