arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13511cs.LGcs.AIcs.CR

大语言模型与经典机器学习在分布偏移和对抗性规避下网络入侵检测的三轴压力测试

A Three-Axis Stress Test of LLM vs Classical ML for Network Intrusion Detection under Distribution Shift and Adversarial Evasion

Muhammad Ebad Atif, Muhammad Haider Ali

AI总结:

本研究通过同数据集、跨数据集和对抗性规避三轴测试,发现XGBoost与RoBERTa-LoRA在网络入侵检测中各有胜负,强调需多轴多指标评估模型。

AI中文摘要:

大型语言模型越来越多地与经典机器学习在网络入侵检测(NIDS)上进行基准测试,几乎总是使用同数据集评估,而该协议被证明是不完整的。在两个独立收集的NetFlow v2网络上,对XGBoost和RoBERTa-LoRA进行三轴评估(同数据集性能、跨数据集迁移和对抗性规避),结果显示没有普遍胜者。两个模型在同数据集上统计上持平。XGBoost在跨数据集分布偏移下决定性胜出,F1分数高出15个百分点,平衡准确率高出25个百分点;在目标网络上,RoBERTa-LoRA的假阳性率达到0.78,尽管F1分数表面适中,但其性能仅略高于随机水平。RoBERTa-LoRA在对抗性规避下决定性胜出,在代表性的中等扰动强度下F1分数高出约17个百分点,而两个模型在整个过程中假阳性率均保持在0.01以下。因此,评估者推荐的模型完全取决于测试哪个轴,而非仅凭同数据集准确率。分阶段特征泄漏消融实验非单调地改善了跨数据集迁移,表明泄漏信号分布在特征表示中,而非局限于少数列,且我们两个网络之间的跨数据集迁移具有强烈方向性。这些结果主张沿多个独立鲁棒性轴评估NIDS模型,且每个轴使用多个指标。

英文摘要:

Large language models are increasingly benchmarked against classical machine learning for network intrusion detection (NIDS), almost always using same-dataset evaluation, and that protocol turns out to be incomplete. Evaluating XGBoost and RoBERTa-LoRA on two independently collected NetFlow v2 networks across three axes (same-dataset performance, cross-dataset transfer, and adversarial evasion) reveals no universal winner. The two models are statistically tied same-dataset. XGBoost wins decisively under cross-dataset distribution shift, by 15 points of F1 and 25 points of balanced accuracy; on the target network RoBERTa-LoRA's false positive rate reaches 0.78, leaving it barely above chance despite a superficially moderate F1. RoBERTa-LoRA wins decisively under adversarial evasion, by roughly 17 points of F1 at a representative mid-range perturbation strength, while both models hold false positive rates below 0.01 throughout. The model an evaluator would recommend therefore depends entirely on which axis is tested, not on same-dataset accuracy alone. A staged feature-leakage ablation improves cross-dataset transfer non-monotonically, indicating the leakage signal is distributed across the feature representation rather than confined to a few columns, and cross-dataset transfer between our two networks is strongly directional. These results argue for evaluating NIDS models along multiple independent robustness axes, and with more than one metric per axis.

补充信息

↑