arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04613cs.LG

为何异常检测算法的排名并不如你所想的可靠

Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think

Simon Klüttermann, Jérôme Rutinowski, Frederik Polachowski, Alice Kirchheim

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过分析7种算法在690个数据集上的表现,发现异常检测算法排名高度不稳定,数据集选择和超参数选择是主要影响因素,可靠基准测试需更大更多样的数据集。

中文摘要 AI 辅助

异常检测是一项对安全性要求极高的机器学习问题,应用范围从欺诈检测到网络入侵预防再到工业监控。尽管已有大量异常检测算法被提出,许多新方法都声称达到了最先进的性能,但不少作者是在彼此不一致的基准设置下得出这一结论的。这种可比性的缺失引发了人们对异常检测基准的可复现性和可靠性的担忧。在本研究中,我们探讨了常见的基准选择对算法排名稳定性的影响。我们使用7种具有代表性的异常检测算法和来自OddBench基准套件的690个数据集,分析了在不同数据集选择、评估指标、超参数配置和随机种子下排名的变化情况。为了量化这种影响,我们引入了一种排名不稳定性指标,用于衡量算法排名在不同基准设置下的变异性。我们的结果表明,异常检测中的算法排名高度不稳定,在很多情况下,几乎所有具有竞争力的算法都能在某些基准配置下成为表现最佳的方法。在所研究的因素中,数据集选择和超参数选择对排名不确定性的影响最大,而随机种子和评估指标的影响相对较小。我们还观察到,可靠的基准测试需要比以往工作中常用的数据集集合大得多、多样化得多的数据集。

英文摘要

Anomaly detection is a safety-critical machine learning problem with applications ranging from fraud detection to network intrusion prevention and industrial monitoring. Despite the large number of proposed anomaly detection algorithms, many novel methods claim state-of-the-art performance. However, many authors do so under benchmark settings that are not aligned with one another. This lack of comparability raises concerns regarding the reproducibility and reliability of anomaly detection benchmarks. In this work, we study the impact of common benchmarking choices on the stability of algorithm rankings. Using seven representative anomaly detection algorithms and 690 datasets from the OddBench benchmark suite, we analyze how rankings change under varying dataset selections, evaluation metrics, hyperparameter configurations, and random seeds. To quantify this effect, we introduce a rank instability metric measuring the variability of algorithm rankings across benchmark settings. Our results show that algorithm rankings in anomaly detection are highly unstable. In many cases, almost every competitive algorithm can appear as the best-performing method under some benchmark configuration. Among the studied factors, dataset selection and hyperparameter choice contribute most strongly to ranking uncertainty, while random seeds and evaluation metrics have a comparatively limited impact. We also observe that reliable benchmarking requires substantially larger and more diverse dataset collections than the ones commonly used in prior work.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • TU Dortmund University(多特蒙德工业大学)
  • Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所)
  • iits Consulting GmbH(iits咨询有限公司)
  • Fraunhofer Institute for Material Flow and Logistics IML(弗劳恩霍夫物流与材料流动研究所IML)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑