发表机构
University of Virginia(弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究从第一性原理构建GIDS-Eval框架,统一评估九个基于图的入侵检测系统,发现现有指标掩盖部署缺陷,并引入GIDS-Lite证明架构复杂性并非检测质量的关键驱动因素。
AI 中文摘要
基于图的网络入侵检测系统(GIDS)报告了强大的基准检测指标,但这些指标几乎不能说明其可部署性。我们从第一性原理出发处理这一问题:不是继承每个已发表系统的预处理、窗口化和阈值设定惯例,而是探究受控比较需要什么,并统一施加这些要求。结果是GIDS-Eval,一个评估框架,它将GIDS分解为六个可互换的阶段,并将这些惯例转化为明确的实验变量,从而使报告的性能可以归因于各个阶段而非整个流水线。我们调查了九个代表性的GIDS,在GIDS-Eval中重新实现了其中五个,并在四种数据集上采用一种匹配的协议对它们进行评估。我们识别出九个反复出现的评估缺口,并量化了每个缺口的影响:两条精心构造的边对八个检测器-数据集组合中三个有隐藏内容的组合实现了完全规避;仅快照窗口一项就导致平均精度(AP)平均相对波动38.3%;跨系统对齐预处理使单个检测器的AP最多移动61.8个百分点;我们重放的18个检测器-数据集组合中,没有一个能在事件到达时发出警报。我们引入了GIDS-Lite,一个在同一框架内构建的无编码器对照系统,它在四个数据集中的两个上以最高575倍的更低运行时间排名AP第一。因此,在我们匹配的协议下,当前基准上架构复杂性并不是检测质量的一致驱动因素,但它确实扩大了操作员必须防御的运行时间、校准和攻击面。
英文摘要
Graph-based network intrusion detection systems (GIDS) report strong benchmark detection metrics, but those metrics establish little about deployability. We approach the problem from first principles: rather than inheriting the preprocessing, windowing, and thresholding conventions of each published system, we ask what a controlled comparison requires and impose it uniformly. The result is GIDS-Eval, an evaluation framework that decomposes a GIDS into six interchangeable stages and turns those conventions into explicit experimental variables, so reported performance can be attributed to individual stages instead of whole pipelines. We survey nine representative GIDS, reimplement five of them within GIDS-Eval, and evaluate them on four datasets under one matched protocol. We identify nine recurring evaluation gaps and quantify the impact of each: two crafted edges achieve full evasion against three of the eight detector-dataset pairs with anything to hide; the snapshot window alone accounts for a mean 38.3% relative swing in average precision (AP); aligning preprocessing across systems moves AP by up to 61.8 percentage points for a single detector; and none of the 18 detector-dataset pairs we replay can alert as events arrive. We introduce GIDS-Lite, an encoder-free control built in the same framework, which ranks first by AP on two of the four datasets at up to 575$\times$ lower runtime. Architectural complexity is therefore not a consistent driver of detection quality under our matched protocol on current benchmarks, but it does enlarge the runtime, calibration, and attack surfaces operators must defend.
CommentsFull version of the paper accepted to the ACM Conference on Computer and Communications Security (CCS) 2026