发表机构
EPITA(EPITA学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DIFFINT自编码器,通过可微区间瓶颈实现数值数据的可解释异常检测,在48个ADBench基准上表现最优,是该领先方法簇中唯一的可解释检测器。
AI 中文摘要
基于重构的异常检测器准确率高但不透明:深度自编码器标记样本却不告知从业者哪些特征范围导致其异常。我们提出DIFFINT,一种自编码器,其潜在瓶颈被构建为一组从原始数值数据端到端直接学习的、软的、轴对齐的区间隶属度,无需任何离散化或二值化。每个潜在单元对应特征空间中一个人类可读的超矩形;实例的编码方式是其相对于其他单元落入每个区间的强度,其重构误差即为异常分数。这既保留了可微表示学习的能力,又提供了可检查的内部结构。我们明确了归纳偏置:对于落在所学支撑的每个活跃坐标之外的点(使用Lipschitz约束的解码器),有经过验证的重构误差下界;对于仅少数特征异常的常见情况,有分级且经经验验证的抑制机制;我们提供了闭式无标签重要性,可根据模型已维护的量对每个(单元、特征)对进行排序,将训练好的区间转化为可审计的候选约束,无需任何异常标签。在48个ADBench基准上,在统一的[-1,1]归一化协议下与22个基线对比,DIFFINT在两个指标上均取得整体最佳平均排名(ROC-AUC为4.10,AUPR为4.16);在仅内点检测器中,它明显领先于同类方法,且与最强的含污染数据检测器具有竞争力(参见分层和完整案例分析)。它是统计关联的七个领先方法簇中唯一的可解释检测器。
英文摘要
Reconstruction-based anomaly detectors are accurate but opaque: a deep autoencoder flags a sample without telling a practitioner which feature ranges made it anomalous. We propose DIFFINT, an autoencoder whose latent bottleneck is structured as a set of soft, axis-aligned interval memberships learned end-to-end directly from raw numerical data, without any discretization or binarization. Each latent unit corresponds to a human-readable hyper-rectangle in feature space; an instance is encoded by how strongly it falls inside each interval relative to the other units, and its reconstruction error is the anomaly score. This keeps the power of differentiable representation learning while exposing an inspectable internal structure. We make the inductive bias precise: a certified reconstruction-error lower bound for points that fall outside every active coordinate of the learned support (with a Lipschitz-enforced decoder), and a graded, empirically verified suppression mechanism for the usual case in which only a few features are abnormal; and we provide a closed-form, label-free importance that ranks each (unit, feature) pair from quantities the model already maintains, turning trained intervals into auditable candidate constraints without ever seeing an anomaly label. On 48 ADBench benchmarks against 22 baselines under a common [-1, 1]-normalized protocol, DIFFINT attains the best mean rank overall on both metrics (4.10 on ROC-AUC, 4.16 on AUPR); among inlier-only detectors it leads its regime clearly, and it is competitive with the strongest contaminated-data detectors (see the stratified and complete-case analyses). It is the only interpretable detector in the statistically-tied leading cluster of seven methods.
CommentsAccepted at ICDM 2026