arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基准与评估协议如何影响基于溯源的入侵检测的结论

How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection

Lorenzo Guerra, Thomas Chapuis, Guillaume Duc, Pavlo Mozharovskyi, Van-Tam Nguyen

arXiv 2608.01454首次发表:更新:

发表机构

Ampere Software Technology(安培软件技术公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现基于溯源的入侵检测系统性能结论受基准和评估协议影响,简单白名单性能可媲美部分学习基线,语义信号质量决定架构差异是否凸显,解读架构主张需结合基准与协议。

AI 中文摘要

基于溯源的入侵检测系统(PIDS)常报告出优异性能,但从这些结果中得出的结论对基准选择和评估协议高度敏感。我们通过在满足审计、标注和校准要求的公开数据集上重新评估代表性PIDS来研究这种依赖性,主要聚焦经审计的DARPA TC E3数据集,采用统一协议,包含时间分离的测试期以及仅用于验证的检查点和阈值校准,并探究哪些架构主张得到了实证支持。我们发现,告警成功率与调查效用可能存在巨大差异,因为部分系统能检测到攻击,却未提供足够的进程级上下文以支持取证调查。在四个主要数据集中的三个,基于训练可执行文件名称和路径构建的简单白名单,在关键操作点指标上与所选学习基线相当或更优,这表明它们测得的大部分性能反映的是词汇新颖性,而非更丰富的溯源建模。为解释为何仅部分数据集能凸显架构差异,我们通过特征完整性和字段熵测量语义信号质量。该分析有助于解释为何部分经审计的E3数据集可凸显告警行为,却无法可靠区分模型架构,而Theia结合了最强的语义信号质量,以及参考模型在排名和节点级恢复上的最显著提升。这些结果表明,PIDS中的架构主张应与产生该主张的基准属性和评估协议一同解读。

英文摘要

Provenance-based intrusion detection systems (PIDS) frequently report strong performance, but the conclusions drawn from these results can be highly sensitive to benchmarking choices and evaluation protocols. We investigate this dependency by re-evaluating representative PIDS on public datasets that meet our audit, labeling, and calibration requirements. Focusing primarily on the audited DARPA TC E3 datasets, we apply a unified protocol with temporally separated test periods and validation-only checkpoint selection and threshold calibration, and ask which architectural claims are empirically supported. We find that alerting success and investigation utility can diverge sharply, as several systems surface attacks without providing enough process-level context to support forensic investigation. Across the four primary datasets, a simple allowlist built from executable names and paths observed during training matches or exceeds the selected learned baselines on key operating-point metrics, showing that comparable performance on these metrics is achievable using lexical novelty alone. Quantifying semantic signal quality through feature completeness and field entropy helps explain why several audited E3 datasets support alerting performance without reliably separating model architectures. In contrast, Theia provides the richest semantic signal and shows the clearest improvements in ranking and node-level recovery for our reference model. Overall, these findings reinforce the importance of interpreting architectural claims in PIDS together with the benchmark properties and evaluation protocol that produced them.

CommentsAccepted at NDSS 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑