发表机构
AGH University of Krakow; American University(克拉科夫AGH科技大学; 美国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究持续异常检测,提出从现有表格异常检测数据集设计可重复基准场景的系统框架,能发现候选任务、过滤不合适任务并得出有原则排序,还能从三个数据集提供五个基准场景,涵盖单数据集和多数据集设置。
AI 中文摘要
持续异常检测(CAD)研究模型如何适应不断演变的数据分布,同时保持在先前观察到的模式上的性能。然而,CAD基准严重依赖于任务的定义、过滤、排序和验证方式。在表格领域,任务边界很少给定,任意分割可能会创建不可学习、冗余或过度可转移的任务,从而掩盖真正的持续学习行为。为此,我们从现有的表格异常检测数据集中引入了一个用于可重复基准场景设计的系统框架。该框架发现候选任务,过滤不合适的任务,并得出能展现不同动态的有原则的排序。该框架使我们能够从三个大规模网络安全异常检测数据集中提供五个基准就绪场景,产生单数据集和多数据集CAD设置。
英文摘要
Continual anomaly detection (CAD) studies how models can adapt to evolving data distributions while retaining performance on previously observed regimes. CAD benchmarks, however, depend critically on how tasks are defined, filtered, ordered, and validated. In tabular domains, task boundaries are rarely given, and arbitrary splits can create unlearnable, redundant, or overly transferable tasks that obscure genuine continual-learning behavior. To this end, we introduce a systematic framework for reproducible benchmark scenario design from existing tabular anomaly-detection datasets. The framework discovers candidate tasks, filters unsuitable tasks, and derives principled orderings that expose diverse dynamics. The framework allows us to deliver five benchmark-ready scenarios from three large-scale cybersecurity anomaly detection datasets, yielding both single-dataset and multi-dataset CAD settings.