arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11969stat.MLcs.LG

我们真的解决了吗?时间序列异常检测中后点调整评估指标的独立对抗性压力测试

Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection

Zongye Lyu

首次发表
浏览论文内容

中文总结 AI 辅助

研究时间序列异常检测中后点调整评估指标修复是否有效,通过对12个指标在多个基准上针对无技能分数生成器进行压力测试,发现修复部分有效,给出指标选择决策协议,如优先选基于PR的指标或PA%K等。

中文摘要 AI 辅助

长期以来,点调整(PA)一直是时间序列异常检测(TSAD)中的默认评分协议,但Kim等人(2022年)表明,它会给随机分数授予近乎完美的F1值。该领域转向了替代指标:PA%K、基于范围的精确率/召回率、归属精确率/召回率以及表面下体积(VUS)ROC/PR。但修复是否有效?迄今为止的每一次稳健性检查都是由提议者运行或理论性的;不存在对真实基准的独立、对抗性、与当前最优方法相关的审计。我们进行了这样一次审计。我们在250个序列的UCR异常存档(主要)以及另外五个基准(SMD、SMAP、MSL、NAB、PSM)上,针对平凡和对抗性无技能分数生成器,对12个采用的指标进行压力测试,依据预先注册的与当前最优方法相关的可博弈性标准(如果无技能生成器达到最佳真实检测器分数的>=90%,则该指标在一个序列上被博弈)。修复只是部分有效:归属F1(在99%的序列上被博弈)以及所测试的每个基于ROC的指标(大致都在62 - 64%)被无技能检测器博弈,确实如此(真实检测器在它们博弈的序列上得分0.97 - 1.00),而基于PR的指标和PA%K能够抵抗(14 - 18%)。有一个决定性的、单向的McNemar分裂:VUS - ROC在119个序列上被博弈,而其姊妹指标VUS - PR没有,反之则没有。ROC超过PR的方向在所有另外五个基准上都重复出现且从不反转,但绝对可博弈率依赖于基准,并且没有单一指标在所有六个基准上都安全(甚至VUS - PR在NAB上也被博弈)。我们发布了一个可通过pip安装的压力测试工具包,其中所有指标实现都在冻结提交时提供,以及一个指标选择决策协议:优先选择基于PR的指标或PA%K,将归属F1和ROC - AUC变体视为可博弈的,并针对每个基准进行验证,而不是假定其稳健性。

英文摘要

Point-adjustment (PA), for years the default scoring protocol in time-series anomaly detection (TSAD), was shown by Kim et al. (2022) to award near-perfect F1 to random anomaly scores. The field adopted a suite of replacement metrics (PA%K, range-based precision/recall, affiliation precision/recall, and Volume-Under-the-Surface, VUS, ROC/PR). We ask, independently and adversarially, whether these resist no-skill detectors on real benchmarks, and find the answer turns entirely on one overlooked variable: N, the number of random attempts an adversary reports the best of. Under a single honest run (N=1), not one replacement metric is gameable on any of six benchmarks (UCR, SMD, SMAP, MSL, NAB, PSM): a random detector reaches 90% of the best real detector's score on at most 11% of series for affiliation-F1, 5% for the ROC family, and 2% for the PR-based metrics and PA%K. But under best-of-N reporting, the seed-shopping endemic to ML, the metrics split sharply. affiliation-F1 and every ROC-based metric inflate steeply, affiliation crossing gameable (25% of series) by N=3 and reaching 0.98 at the full pool (N=41), the ROC family crossing by N=9-11; the PR-based metrics and PA%K stay near-flat at every N, floored near the anomaly prevalence (the lone exception is NAB at large N). A paired test finds VUS-ROC inflated on 131 series where its sibling VUS-PR is not, and never the reverse. The ROC-vs-PR split follows from the order-statistic behaviour of AUC under extreme class imbalance (a random PR-AUC is floored at prevalence); affiliation inflates by a second route, its extreme single-run leniency (already fragile at N=1). We release a pip-installable stress-test harness, and recommend reporting single-run scores or disclosing N and preferring PR-based metrics, which resist best-of-N inflation on nearly every benchmark.

发表机构

  • Faculty of Information Technology, Monash University(信息科技学院,墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑