基于表格数据的流匹配无监督异常检测
Unsupervised Anomaly Detection Using Flow Matching on Tabular Data
浏览论文内容
中文总结 AI 辅助
本研究针对受污染训练数据下的表格异常检测,对比TCCM与Forest-Flow,发现轨迹类异常评分更稳定,Forest-Flow表现可与TCCM媲美甚至更优,凸显异常评分的重要性。
中文摘要 AI 辅助
金融异常检测通常依赖大量未标记的交易日志,其中异常样本可能在训练期间已存在。这种训练集污染违背了许多异常检测方法所基于的干净正常数据假设。尽管流匹配在生成式建模中表现出强大性能,但它在无监督表格异常检测中的鲁棒性仍未得到充分探索。本研究针对受污染训练数据下基于流匹配的异常检测展开,将时间条件收缩匹配(TCCM)与森林流(Forest-Flow)进行对比,并评估多种异常评分函数。结果表明,异常评分的选择至关重要:TCCM 所用的原始单步决策评分对污染敏感,而基于轨迹的偏差评分与重建评分能提供更稳定的异常信号。采用这些评分后,森林流的表现可与 TCCM 相媲美,部分情况下甚至优于 TCCM。这些发现凸显了在严重类别不平衡的金融异常检测场景中,异常评分对流匹配方法的重要性。
英文摘要
Financial anomaly detection often relies on large unlabeled transaction logs, where anomalous samples may already be present during training. Such training-set contamination violates the clean-normal data assumption underlying many anomaly detection methods. Although flow matching has demonstrated strong performance in generative modeling, its robustness in unsupervised tabular anomaly detection remains underexplored. In this work, we study flow-matching-based anomaly detection under contaminated training data by comparing Time-Conditioned Contraction Matching (TCCM) with Forest-Flow and evaluating multiple anomaly scoring functions. Our results show that the choice of anomaly score is critical. The original single-step Decision score used by TCCM is sensitive to contamination, whereas trajectory-based Deviation and Reconstruction scores provide more stable anomaly signals. With these scores, Forest-Flow becomes competitive with, and in some cases outperforms, TCCM. These findings highlight the importance of anomaly scoring for flow-matching methods in financial anomaly detection under severe class imbalance.
发表机构
- University of Mannheim(曼海姆大学)
- MPI for Informatics(马克斯·普朗克信息学研究所)
- Saarland Informatics Campus(萨尔兰信息学园区)
机构由 AI 辅助整理,请以论文原文为准。