arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

检测、解释、解读:时间序列异常检测、可解释性与可解读性的端到端基准

Detect, Explain, Interpret: An End-to-End Benchmark for Time Series Anomaly Detection, Explainability and Interpretability

Roberto Stanzione, Jules Barbe, Magali Parrino, Jérémie Fourmann, Paul Boniol

arXiv 2610.01168首次发表:更新:

发表机构

Inria; ENS; CNRS; PSL; Scality; EDF(法国国家信息与自动化研究所; 巴黎高等师范学院; 法国国家科学研究中心; 巴黎文理研究大学; Scality公司; 法国电力集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有时间序列异常检测基准忽视可解释性与可解读性的问题,提出完全注释的SHAD基准,含215个真实云存储系统的高维时间序列,并评估检测、可解释性及可解读性基线方法。

AI 中文摘要

时间序列异常检测因复杂时间序列数据的日益普及而受到越来越多的关注。这一热潮催生了众多检测方法的发展,以及旨在全面评估其性能的各种基准。然而,大多数现有检测器在很大程度上仍对领域背景不敏感,忽视了可解释性和可解读性。造成这一差距的主要原因之一是,当前基准主要关注检测准确性,只有少数基准评估了空间可解释性。此外,目前没有任何基准提供足够丰富的语义注释来支持生成人类可理解的异常解读。为解决这些局限,我们引入了SHAD(Scality高维异常检测基准),这是一个完全注释的基准,由Scality运营的真实世界分布式云存储系统收集的215个多变量、高维时间序列组成。所提出的数据集包含丰富的上下文信息,涵盖三类严重程度各异的异常。作为进一步贡献,我们通过评估检测、可解释性和可解读性的基线方法,为未来工作提供了基础,覆盖了TSAD流水线的所有阶段。对于检测,我们对广泛的现有异常检测器进行了基准测试,测试它们在所提出的真实世界数据集上的有效性。然后,我们考虑可解释性,评估测量生成异常分数中每个维度的贡献是否能提供准确的异常归因。最后,对于可解读性,我们研究了冻结LLM基线在定位和解读异常方面的有效性。

英文摘要

Time Series Anomaly Detection has received increasing attention, driven by the growing availability of complex time series data. This surge has led to the development of numerous detection methods, as well as a variety of benchmarks aimed at thoroughly evaluating their performance. However, most existing detectors remain largely agnostic to domain context, overlooking explainability and interpretability. One of the main reasons for this gap is that current benchmarks primarily focus on detection accuracy, and only few of them evaluate spatial explainability. Moreover, no benchmark currently provides sufficiently rich semantic annotations to support the generation of human-understandable interpretations of anomalies. To address these limitations, we introduce SHAD (Scality High-dimensional Anomaly Detection benchmark), a fully annotated benchmark composed of 215 multivariate, high-dimensional time series collected from real-world distributed cloud storage systems operated by Scality. The proposed dataset includes rich contextual information, covering three families of anomalies with varying degrees of severity. As further contribution, we provide a foundation for future work by evaluating baseline methods for Detection, Explainability, and Interpretability, covering all stages of a TSAD pipeline. For Detection, we benchmark a wide range of existing anomaly detectors, testing their effectiveness on the proposed real-world dataset. Then, we consider explainability by evaluating whether measuring the contribution of each dimension in the generated anomaly score can provide accurate anomaly attributions. Finally, for interpretability, we investigate the effectiveness of frozen LLM baselines in localizing and interpreting anomalies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑