超越检测准确率:面向资源感知的物联网入侵检测的解释成本、稳定性与效用测量
Beyond Detection Accuracy: Measuring Explanation Cost, Stability, and Utility for Resource-Aware IoT Intrusion Detection
- organization= Independent Researcher , city= Istanbul , country= Türkiye
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对物联网入侵检测,构建泄露安全的CICIoT2023语料库,评估四种模型的预测性能、解释成本与稳定性,发现需综合多指标而非仅检测准确率实现实用可解释检测。
AI中文摘要:
机器学习入侵检测研究通常侧重预测准确率,却将解释生成视为计算上免费的后处理步骤。本研究针对二进制物联网(IoT)入侵检测,联合评估预测有效性、解释成本、局部解释稳定性及选择性解释。构建了泄露安全的CICIoT2023语料库,采用精确39特征哈希、非有限值处理、精确特征去重、保守标签碰撞移除及确定性哈希级划分。在自然分布与平衡分布测试集上评估了逻辑回归(Logistic Regression)、决策树(Decision Tree)、随机森林(Random Forest)和XGBoost模型。测量了TreeSHAP的成本,在保留预测的扰动下评估稳定性,并采用验证校准策略分配解释工作量。XGBoost提供最强的整体预测性能,随机森林的误报率最低。在5000个样本下,TreeSHAP对随机森林需700.759秒,对XGBoost仅需1.471秒。随机森林表现出最强的基础级解释稳定性;XGBoost保持较高的特征排序和方向一致性,但存在更明显的顶部特征更替及归因幅度漂移。在平衡测试中,约90%的漏报解释覆盖可实现28%-32%的计算节省,约95%覆盖可实现15%-23%的节省;在攻击占比高的自然流行分布下,节省幅度小得多。这些结果表明,可操作的实用可解释物联网入侵检测依赖于预测质量、解释成本、局部稳定性、工作量流行度及选择性调用,而非仅检测准确率。
英文摘要:
Machine-learning intrusion-detection studies commonly emphasize predictive accuracy while treating explanation generation as a computationally free post-processing step. This study jointly evaluates predictive effectiveness, explanation cost, local explanation stability, and selective explanation for binary Internet of Things (IoT) intrusion detection. A leakage-safe CICIoT2023 corpus was constructed using exact 39-feature hashes, non-finite-value handling, exact-feature deduplication, conservative label-collision removal, and deterministic hash-level partitioning. Logistic Regression, Decision Tree, Random Forest, and XGBoost were evaluated on natural and balanced test distributions. TreeSHAP cost was measured, stability was assessed under prediction-preserving perturbations, and validation-calibrated policies were used to allocate explanation workload. XGBoost provided the strongest overall predictive profile, while Random Forest produced the lowest false-positive rate. At 5,000 samples, TreeSHAP required 700.759 s for Random Forest and 1.471 s for XGBoost. Random Forest showed the strongest overall base-level explanation stability; XGBoost retained high rank and directional consistency but showed greater top-feature turnover and attribution-magnitude drift. On the balanced test, about 90% false-negative explanation coverage permitted 28-32% compute savings, while about 95% coverage permitted 15-23% savings. Savings were much smaller under the attack-heavy natural prevalence. These results show that operationally useful explainable IoT intrusion detection depends on predictive quality, explanation cost, local stability, workload prevalence, and selective invocation rather than detection accuracy alone.