arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28725cs.CRcs.LGcs.NI

揭开物联网入侵检测中的捷径学习:特征依赖与数据泄漏的法证式多范式评估

Unmasking Shortcut Learning in IoT Intrusion Detection: A Forensic, Multi-Paradigm Evaluation of Feature Dependence and Data Leakage

Uday Shankar Roy, Mahbuba Jahan Minu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过多范式评估揭示物联网入侵检测模型存在捷径学习,证明性能受特征表示限制,时间戳可被非线性模型利用,并提出4点评估协议清单以促进真实泛化。

中文摘要 AI 辅助

基于机器学习的网络入侵检测系统在物联网基准测试中常报告近乎完美的性能。然而,这些模型是学习了可泛化的攻击行为,还是利用了虚假的数据集捷径(如静态测试平台的IP/MAC地址和按时间顺序记录的伪影),仍是一个重要问题。我们评估了CyberFlowIoT-GICAP基准,其中包含126个PCAP会话中的3,617,388条流记录,以及849,395条良性流。我们使用PCAP不相交划分,在四种特征配置下评估了四种学习范式;此外,还使用传统的随机流划分对LightGBM进行了评估。当仅使用统计流行为(Fbehav)时,LightGBM(92.58%±8.18%)、随机森林(92.59%±8.18%)和深度MLP(92.55%±8.18%)实现了几乎相同的Macro-F1,表明性能受特征表示而非模型复杂度的限制。使用原始时间戳(Ftstamp)时,基于树的模型达到99.28%的Macro-F1,而线性模型仍为90.62%,表明非线性模型可以利用数据集特定的时间结构。攻击可检测性高度不对称:在非线性模型中,高速率和主动攻击仅凭流行为即可保持>99.8%的召回率,而DNS信标在去除上下文特征后,召回率从27.78%降至0.00%。传统的随机流划分将攻击召回率最多提高14.00%,凸显了将同一会话的流同时放入训练集和测试集的影响。最后,我们提出了一个包含4点的协议检查清单,用于现实物联网NIDS评估。

英文摘要

Machine learning-based Network Intrusion Detection Systems often report near-perfect performance on IoT benchmarks. However, whether these models learn generalizable attack behavior or exploit spurious dataset shortcuts- such as static testbed IP/MAC addresses and chronological recording artifacts-remains an important question. We evaluate the CyberFlowIoT-GICAP benchmark, containing 3,617,388 flow records across 126 PCAP sessions with 849,395 benign flows. Four learning paradigms are evaluated across four feature configurations using PCAP-disjoint splits; LightGBM is additionally evaluated using conventional random-flow splitting. When only statistical flow behavior is used (Fbehav), LightGBM (92.58% +/- 8.18%), Random Forest (92.59% +/- 8.18%), and Deep MLP (92.55% +/- 8.18%) achieve nearly identical Macro-F1, indicating that performance is constrained by feature representation rather than model complexity. With raw timestamps (Ftstamp), tree-based models reach 99.28% Macro-F1, while the linear model remains at 90.62%, showing that nonlinear models can exploit dataset-specific temporal structure. Attack detectability is highly asymmetric: high-rate and active attacks maintain >99.8% recall from flow behavior alone in nonlinear models, whereas the DNS Beaconing drops from 27.78% to 0.00% recall when contextual features are removed. Conventional random-flow splitting increases attack recall by up to 14.00%, highlighting the effect of placing flows from the same sessions in both training and test sets. We conclude with a 4-point protocol checklist for realistic IoT NIDS evaluation.

发表机构

  • Khulna University of Engineering & Technology(库尔纳工程技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑