发表机构
Saarland University; German Research Center for Artificial Intelligence (DFKI); Leipzig University(萨尔大学; 德国人工智能研究中心(DFKI); 莱比锡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对日记账分录异常检测,提出类型感知评估方法(含FSR指标),发现高命中率可掩盖盲点,反馈强化模式但不扩大覆盖。
AI 中文摘要
日记账分录异常检测器通常使用ROC-AUC、精确率和召回率在整个总体上进行评估,忽略了审查预算以及发现了哪些异常类型。我们提出了一种类型感知评估方法,结合了每类型召回率、公平份额类型召回率(FSR)(该指标将每种类型的信用上限设为其预算份额)、类型覆盖率和首次命中排名。我们在四个带有注入类型异常的真实客户分类账和一个公开的合成分类账上,评估了九种无监督检测器、一种有监督参考方法以及反馈驱动的深度半监督异常检测(DeepSAD)。在最大的客户分类账上,主成分分析(PCA)、自编码器(AE)和变分自编码器(VAE)各自在前100条分录中平均放置了98个异常,但其中至少有95.8个属于同一类型。FSR反而青睐最近邻(kNN)检测器,并在四个客户分类账中的三个上改变了排名最高的检测器。表示方式也很重要:独热编码暴露了未见过的账户,而频率编码在很大程度上未能检测到未见过的对应账户。在公开分类账上,基于直方图的离群值分数(HBOS)和基于经验累积分布的离群值检测(ECOD)在1,386条分录内达到了所有八个标记,而kNN(在1,000条分录时是命中领先者)首次在排名4,641处达到交叉链接清算,而有监督的行级参考方法在1,000条分录内未命中该标记。在那里,自适应DeepSAD审查协议将每100次审查的平均命中数从40.0提高到68.3,但类型覆盖率仅从2.7提高到3.0。这些发现表明,高命中率可能掩盖系统性的盲点,并表明反馈可以强化现有的检测模式,而不会扩大异常覆盖范围。
英文摘要
Journal entry anomaly detectors are commonly evaluated on the full population with ROC-AUC, precision and recall, ignoring the review budget and which anomaly types are found. We propose a type-aware evaluation combining per-type recall, fair-share type recall (FSR), which caps each type's credit at its budget share, type coverage and first-hit rank. We evaluate nine unsupervised detectors, a supervised reference and feedback-driven Deep Semi-Supervised Anomaly Detection (DeepSAD) on four real client ledgers with injected typed anomalies and a public synthetic ledger. On the largest client ledger, principal component analysis (PCA), an autoencoder (AE) and a variational autoencoder (VAE) each place on average 98 anomalies among the first 100 postings, but at least 95.8 belong to one type. FSR instead favours a nearest-neighbour (kNN) detector and changes the top-ranked detector on three of four client ledgers. Representation also matters: one-hot encoding exposes unseen accounts, whereas frequency encoding leaves unseen contra accounts largely undetected. On the public ledger, the Histogram-Based Outlier Score (HBOS) and Empirical Cumulative Distribution-Based Outlier Detection (ECOD) reach all eight markings within 1,386 entries, whereas kNN, the hit leader at 1,000 entries, first reaches cross-linked clearing at rank 4,641, and the supervised row-level reference misses this marking within 1,000 entries. There, the adaptive DeepSAD review protocol raises mean hits per 100 reviews from 40.0 to 68.3 but type coverage only from 2.7 to 3.0. These findings show that high hit rates can conceal systematic blind spots and suggest that feedback can reinforce existing detection patterns without broadening anomaly coverage.