发表机构
Middlesex University; Humboldt-Universität zu Berlin; Shahid Beheshti University; Kar Higher Education; Tabriz University of Medical Sciences(密德萨斯大学; 柏林洪堡大学; 沙希德贝赫什提大学; 卡拉高等教育学院; 大不里士医科大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对反洗钱检测中规则方法误报率高的问题,提出基于图神经网络的混合框架,联合优化预测性能、解释忠实性与不确定性校准,在HI-Small基准上显著优于基线模型。
AI 中文摘要
目的:洗钱活动威胁金融体系,而基于规则的监控存在高误报率且难以捕捉交易关系模式的局限。本研究提出一种可解释的基于图结构的反洗钱(AML)检测框架,同时解决预测性能、解释忠实性和不确定性校准问题。方法:使用IBM反洗钱交易(HI-Small)基准数据集,该数据集包含518,573个账户间的5,078,345笔交易,其中非法交易占比0.10%,通过时间安全、防泄漏的划分方式构建了一个包含4,487,133条边的有向属性图。将三种GATv2架构(BASE、BASE-Large和IMPROVED,其中IMPROVED融合了双向消息传递、边更新和端口感知特征)与使用相同特征的XGBoost和随机森林进行比较。最佳模型使用GNNExplainer对照文档化的洗钱类型模式进行评估,并使用Mondrian共形预测进行不确定性校准。结果:IMPROVED模型取得了最强性能(AUPRC = 64.85%,最佳F1 = 68.42%),比BASE-Large高出30.65个AUPRC点,比XGBoost(AUPRC = 38.46%)高出26.4个点。所提出的机制带来的提升超过了单纯增加模型容量。GNNExplainer恢复文档化类型模式的忠实度高于注意力权重基线和随机基线(平均Jaccard重叠:21.4% 对比 1.0%,p < 0.001)。Mondrian共形预测达到了接近90%名义目标的覆盖率,平均预测集大小接近1。结论:显式的交易拓扑建模显著提升了AML检测性能,而忠实的解释和校准的不确定性支持可解释的、人在回路的合规决策。
英文摘要
Purpose: Money laundering threatens financial systems, while rule-based monitoring suffers from high false-positive rates and limited ability to capture relational transaction patterns. This study proposes an explainable graph-based framework for anti-money laundering (AML) detection that jointly addresses predictive performance, explanation faithfulness, and uncertainty calibration. Methods: Using the IBM Transactions for Anti-Money Laundering (HI-Small) benchmark, comprising 5,078,345 transactions among 518,573 accounts with 0.10% illicit transactions, a directed attributed graph with 4,487,133 edges was constructed using temporal, leakage-safe partitioning. Three GATv2 architectures, BASE, BASE-Large, and IMPROVED, incorporating bidirectional message passing, edge updates, and port-aware features, were compared with XGBoost and Random Forest using identical features. The best model was evaluated using GNNExplainer against documented typologies and Mondrian conformal prediction for uncertainty calibration. Results: IMPROVED achieved the strongest performance (AUPRC = 64.85%, Best F1 = 68.42%), exceeding BASE-Large by 30.65 AUPRC points and XGBoost (AUPRC = 38.46%) by 26.4 points. The proposed mechanisms contributed more than capacity scaling alone. GNNExplainer recovered documented typologies with higher fidelity than attention-weight and random baselines (mean Jaccard overlap: 21.4% vs. 1.0%, p < 0.001). Mondrian conformal prediction achieved coverage close to the 90% nominal target with an average prediction-set size near one. Conclusion: Explicit transaction topology modeling substantially improves AML detection, while faithful explanations and calibrated uncertainty support interpretable, human-in-the-loop compliance decision-making.