发表机构
INFOCZ Inc; Seoul National University(INFOCZ公司; 首尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究审计九个基于溯源入侵检测实现,揭示三种测量效应(标签选择、检查点选择、缓冲区缺陷),质疑警报精确率声明,并提出六项可核查报告要求。
AI 中文摘要
复现基于溯源的入侵检测器的评分,并不能确定该评分对其发出的警报或其编码器使用的信息有何说明。我们审计了九个已发布的实现,使用其自身代码执行了四个检测器,并分离出三种测量效应。第一,固定警报比较将标签选择与邻域信用分开:ThreaTrace在采用邻域信用时报告精确率为0.938,尽管其994个警报中仅有7个带有其自身的攻击标签。第二,移除测试标签检查点选择会使攻击检测精确率(ADP)在四个四十成员word2vec配置中降低0.16至0.33,且不改变检测器顺序。第三,一个缓冲区重用缺陷使线性编码器获得非预期的依赖度输入。修正该缺陷会在两台主机上的每个相同输入初始化对中降低仅类型ADP,而历史word2vec效应则取决于主机。这些发现对受检实现中关于警报精确率、性能幅度和静态属性充分性的声明提出了质疑。STRICT将这些发现与六项可核查的报告要求联系起来。由于比较以基准目标为条件,它们既未验证这些标签,也未确立通用的检测器排名。
英文摘要
Reproducing a provenance-based intrusion detector's score does not establish what that score says about its emitted alarms or the information its encoder uses. We audit nine released implementations, execute four detectors using their own code, and isolate three measurement effects. First, a fixed-alert comparison separates label choice from neighbourhood credit: ThreaTrace reports precision 0.938 with neighbourhood credit, although only seven of its 994 alarms carry its own attack label. Second, removing test-label checkpoint selection lowers attack detection precision (ADP) by 0.16 to 0.33 across four forty-member word2vec configurations without changing detector order. Third, a buffer-reuse defect gives a linear encoder unintended degree-dependent inputs. Correcting it lowers type-only ADP in every identical-input initialization pair on two hosts, while historical word2vec effects depend on the host. These findings qualify claims from the inspected implementations about alarm precision, performance magnitude and static-attribute sufficiency. STRICT connects them to six checkable reporting requirements. Because the comparisons condition on benchmark targets, they neither validate those labels nor establish a universal detector ranking.