AI 中文总结
该研究提出一种校准型可解释双峰机器学习框架,通过特征提取、混合重采样、保序回归校准等技术,在CIC-IDS2017数据集上实现了高精度的混合入侵检测,可检测未知攻击且决策可解释。
AI 中文摘要
现代通信系统面临着关键的检测缺口,由于极端的数据不平衡和黑箱决策逻辑,难以检测未知攻击和稀有威胁类别。我们提出了一种面向网络安全的校准型可解释双峰机器学习(ML)框架,该框架在不使用深度学习的复杂性的情况下,将已知类别的精度与开集泛化能力统一起来。我们的框架引入了面向安全的特征提取以提高信噪比,采用混合重采样(ADASYN + 手动提升)以减少类别不平衡,使用保序回归校准和自适应阈值(XSS的阈值τ=0.30)以恢复稀有攻击的召回率,并采用基于SHAP的可解释性来验证与领域对齐的决策逻辑。在CIC-IDS2017数据集上进行评估,并与先前的ML模型和研究进行比较,我们的框架在已知攻击上取得了显著的准确率(Macro F1=0.8626),且在1%的FPR下检测未知类别的TPR最高可达90.17%(DoS slowloris)和77.04%(Web-XSS)。SHAP分析证实,决策由与安全相关的特征驱动,而非模型的人工产物。我们的工作通过在单一可复现的框架中提供校准、可解释且具备开集能力的攻击检测与预防,弥合了理论模型与操作型入侵检测系统(IDS)之间的差距。
英文摘要
Modern communication systems face critical gaps in detecting unknown attacks and rare threat classes due to extreme data imbalance and black-box decision logic. We propose a bimodal framework of calibrated and explainable machine learning (ML) for network security, unifying known-class precision with open-set generalization without the complexity of deep learning. Our framework introduces security-oriented feature extraction to enhance signal-to-noise ratio, hybrid resampling (ADASYN + manual boosting) to reduce class imbalance, isotonic calibration and adaptive thresholding ($τ=0.30$ for XSS) to recover recall for rare attacks, and SHAP-based explainability to validate domain-aligned decision logic. Evaluated on the CIC-IDS2017 dataset and compared with prior ML models and studies, our framework achieves significant accuracy on known attacks (Macro F1 = 0.8626) and detects unknown classes at 1% FPR with TPR up to 90.17% (DoS slowloris), and 77.04% (Web-XSS). The SHAP analysis confirms decisions are driven by security-relevant features, not model artifacts. Our work bridges the gap between theoretical models and operational IDS by delivering calibrated, explainable, and open-set-capable attack detection and prevention in a single, reproducible framework. Keywords: intrusion detection, cybersecurity and privacy, explainable AI, machine learning