AI 中文总结
CORE通过去相关特征对齐模块和上下文内重构方法,解决统一表格异常检测中异构数据对齐及二分类任务的局限,可支持任意未见数据集的统一异常检测。
AI 中文摘要
表格异常检测(TAD)专注于识别表格数据中偏离多数的异常样本,已受到越来越多关注。近年来,统一TAD成为新兴趋势,旨在用单个可泛化模型检测不同数据集的异常。在统一TAD中,对齐异构数据仍具挑战性:现有方法常依赖基于距离的统一特征构建,可能掩盖原始特征的语义;且现有方法通常将异常检测表述为二分类任务,可能忽略不同数据集的多样异常模式,还可能被无代表性的合成异常误导。为应对这些挑战,我们提出统一TAD的上下文内重构方法(简称CORE)。它引入去相关特征对齐模块,将异构特征直接对齐到保留语义信息的统一表示空间;同时将统一TAD表述为上下文内重构问题,无需标记或合成异常。具体而言,上下文内重构模块利用上下文正常样本重构每个样本,以捕获特定数据集的分布,使重构误差反映样本与常态的偏差,从而支持对任意未见数据集的统一TAD。
英文摘要
Tabular anomaly detection (TAD), which focuses on identifying abnormal samples that deviate from the majority in tabular data, has received growing attention. Recently, there has been an emerging trend towards unified TAD, which seeks to detect anomalies across different datasets using a single generalizable model. In unified TAD, aligning heterogeneous data remains challenging. While existing methods often rely on distance-based unified feature construction, they may obscure the semantics of the original features. Moreover, existing approaches typically formulate anomaly detection as a binary classification task, which may overlook diverse anomaly patterns from various datasets and be misled by unrepresentative synthetic anomalies. To address these challenges, we propose an in-COntext REconstruction approach for unified TAD (CORE for short). It introduces a decorrelated feature alignment module to directly align heterogeneous features into a unified representation space, which retains their semantic information. Meanwhile, CORE formulates unified TAD as an in-context reconstruction problem, eliminating the need for labeled or synthesized anomalies. Specifically, the in-context reconstruction module reconstructs each sample by leveraging contextual normal samples to capture dataset-specific distributions, such that reconstruction errors reflect its deviation from normality, facilitating unified TAD on arbitrary unseen datasets.