先击败计数器:时序图异常检测器的基线
Beat the Counter First: A Baseline for Temporal-Graph Anomaly Detectors
浏览论文内容
中文总结 AI 辅助
该研究针对时序图异常检测,提出无参数的SimpleCount单特征基准,经多数据集实验表明其性能优于部分复杂模型且计算成本低,建议异常检测增益需结合计算成本与单特征基准评估。
中文摘要 AI 辅助
流式边级图异常检测(GAD)的进展以日益复杂的架构为标志,从计数-最小草图卡方检验到内存增强型注意力网络。然而,这种额外复杂性带来的经验增益尚未得到系统评估。我们提出SimpleCount,这是一种无参数拟合的基准方法,从计数、时效性、首次出现指标及计数衍生变换构成的固定池中为每个数据集选择一个标量特征。我们在5个公开数据集和1个合成数据集上将SimpleCount与两个时序图检测器模型及拟合完整特征向量的IsoForest对照进行比较。SimpleCount在6个数据集中的3个上与SLADE表现相当或更优,在全部6个数据集上优于IsoForest。我们报告了配对统计检验和5次随机种子的SLADE评估结果:SLADE的挂钟时间比SimpleCount多23至133倍。在Synth-Triangle及额外的Synth-Quad探测中,事件前结构分数以高达0.955的AUC恢复了植入信号,而所有被评估的检测器模型均接近随机水平。复杂性的益处取决于数据集,每项宣称的增益都应结合其计算成本与强大的单特征基准进行报告。
英文摘要
Progress in streaming, edge-level graph anomaly detection (GAD) has been marked by increasingly elaborate architectures, from count-min-sketch chi square tests to memory-augmented attention networks. Yet the empirical gains attributable to this added complexity have not been systematically evaluated. We propose SimpleCount, a reference with no parameter fitting that selects one scalar feature per dataset from a fixed pool of counts, recencies, first-occurrence indicators, and count-derived transforms. We compare SimpleCount with two temporal-graph detector models and an IsoForest control fitted to the complete feature vector across five public datasets and one synthetic dataset. SimpleCount matches or exceeds SLADE on three of six datasets and exceeds IsoForest on all six. We report paired statistical tests and five-seed SLADE evaluations. SLADE requires 23 to 133x more wall-clock time than SimpleCount. On Synth-Triangle and an additional Synth-Quad probe, pre-event structural scores recover the planted signal at AUC up to 0.955, while all evaluated detector models remain near random. The benefit of complexity is dataset-dependent, and every claimed gain should be reported against a strong one-feature reference together with its compute cost.