适应先验数据拟合网络用于表格异常检测
Adapting prior-data fitted networks for tabular anomaly detection
浏览论文内容
中文总结 AI 辅助
本研究利用先验数据拟合网络(PFN)的特征进行表格异常检测,通过冻结特征(ZEN)和微调(FOCUS)方法在ADBench基准上取得最优性能。
中文摘要 AI 辅助
虽然深度特征已经改变了图像和视频中的异常检测,但它们对表格数据的影响却没那么显著,部分原因是缺乏强大的深度表示。最近,先验数据拟合网络(PFNs)已成为表格数据此类表示的一个有前景的来源。在这项工作中,我们研究了如何调整和利用PFN表示进行异常检测。这个问题比看起来更难。部署前没有可用的异常样本,因此模型参数无法通过监督进行调整,而且定义正常行为的参考集本身可能包含它本应揭示的异常。我们首先使用冻结的TabPFN特征开始研究。通过每个样本到其特征空间中最近邻的距离进行评分,已经取得了很强的结果。我们确定了要使用的层和适合该任务的特征提取过程。接下来,为了进一步提高性能,我们使用参考集对模型进行微调,使得所得特征能更好地区分正常样本和异常样本。在ADBench基准上,我们无需微调的方法(ZEN)达到了比所有基线更高的平均AUROC,而我们的微调方法(FOCUS)进一步提高了性能。我们的方法也适用于各种PFN模型。
英文摘要
While deep features have transformed anomaly detection in images and video, their impact on tabular data has been less substantial, partly due to the limited availability of strong deep representations. Recently, prior-data fitted networks (PFNs) have emerged as a promising source of such representations for tabular data. In this work, we investigate how PFN representations can be adapted and leveraged for anomaly detection. The question is harder than it looks. No anomalies are available before deploy- ment, so model parameters cannot be tuned with supervision, and the reference set that defines normal behavior may itself contain the very anomalies it is supposed to reveal. We begin our study using frozen TabPFN features. Scoring each sam- ple by its distance to its nearest neighbors in feature space already gives strong results. We identify which layers to use and a feature-extraction procedure suited to the task. Next, to further improve performance, we use the reference set to fine- tune the model, so that the resulting features better separate normal samples from anomalies. On the ADBench benchmark, our fine-tuning free approach (ZEN) reaches a higher mean AUROC than every baseline, and our fine-tuned method (FOCUS) improves on it further. Our approach also generalizes across PFN models.
发表机构
- Technion – Israel Institute of Technology(以色列理工学院)
机构由 AI 辅助整理,请以论文原文为准。