arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

选择规则决定胜负:开放集图异常检测的预注册审计

The Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly Detection

Farhan Shahriyar Hossain, Taufikur Rahman Fuad, Md Abrar Jahin, Md Rizwan Parvez

arXiv 2609.33370首次发表:更新:

发表机构

Islamic University of Technology; University of Southern California; Qatar Computing Research Institute(伊斯兰技术大学; 南加州大学; 卡塔尔计算研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过预注册审计发现,测试得分读取规则(最佳epoch vs. 验证)显著影响开放集图异常检测方法的排名,且伪标签仅在重标注类图上有效,并提供了报告清单。

AI 中文摘要

开放集图异常检测在仅使用来自一个类别的少量标注异常样本进行训练的同时,还必须发现从未被标注过的异常类别。已发表的研究结果共享三个惯例:测试得分在测试集上的最佳epoch处读取,基线数值从早期论文中复制,且大多数异常是重新标注为异常的少数类。我们探究这些惯例在多大程度上决定了报告的排名。我们重新运行了两种近期方法DEMO和NSReg,以及OUTPOST——一个为本研究构建的小型一阶检测器。所有三种方法在八个图(基线方法无法在ogbn-mag上运行,因此为七个)上使用相同的协议,具有相同的随机种子和划分,每种方法运行十次,每次运行均在最佳epoch规则和可部署的验证规则下进行评分。在运行测试之前,我们注册了40个预测。三项发现成立。第一,规则改变了领先者:在最佳epoch规则下,OUTPOST和NSReg各自在七个图中的三个上领先,而在验证规则下,NSReg在五个上领先。第二,最佳epoch的增益取决于基准的构建方式:在三个小的重新标注类图上为0.045--0.080 AUC-ROC,在三个真实欺诈图上为0.002--0.014。第三,OUTPOST中的伪标签在相同的三个图上价值0.038--0.065 AUC-ROC,但在任何真实欺诈图上均无益处。我们还表明,用于超参数选择的0.002并列带在所有六个测试图上低于配对标准误差,即使在十个种子下也是如此。我们的40个预测中有12个被证伪,我们对此进行了报告。最后,我们提供了一个简短的报告清单。

英文摘要

Open-set graph anomaly detection trains on a few labeled anomalies from one class and must also find anomaly classes that were never labeled. Published results share three conventions: the test score is read at the best epoch on the test set, baseline numbers are copied from earlier papers, and most anomalies are minority classes relabeled as anomalous. We ask how much of the reported ranking these conventions decide. We re-run two recent methods, DEMO and NSReg, together with OUTPOST, a small first-order detector built for this study. All three use one protocol with identical seeds and splits on eight graphs (seven for the baselines, which cannot run on ogbn-mag), ten seeds each, and every run is scored under both the best-epoch rule and a deployable validation rule. Before the runs that test them, we registered 40 predictions. Three findings hold. First, the rule changes the leader: under the best-epoch rule, OUTPOST and NSReg each lead three of seven graphs, while under the validation rule, NSReg leads five. Second, the best-epoch bonus depends on how the benchmark was built: 0.045--0.080 AUC-ROC on the three small relabeled-class graphs and 0.002--0.014 on the three real fraud graphs. Third, pseudo-labeling in OUTPOST is worth 0.038--0.065 AUC-ROC on the same three graphs but gives no benefit on any real fraud graph. We also show that a 0.002 tie band for hyperparameter selection lies below the paired standard error on all six graphs tested, even at ten seeds. Twelve of our 40 predictions were falsified, and we report them. We close with a short reporting checklist.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑