KONTOGRAPH:200毫秒决策预算下实时反洗钱的经验证时点特征一致性与摊销解释
KONTOGRAPH: Verified Point-in-Time Feature Consistency and Amortised Explanation for Real-Time Anti-Money Laundering under a 200 ms Decision Budget
浏览论文内容
中文总结 AI 辅助
该研究针对欧盟2024/886号条例的10秒结算要求,构建了200毫秒决策预算的KONTOGRAPH反洗钱流水线,发现时间图网络性能显著提升、特征编译存在时点违规、ONNX转换会影响决策。
中文摘要 AI 辅助
欧盟《2024/886号条例》要求欧洲支付服务提供商全天候在10秒内完成欧元转账结算,这取消了传统反洗钱(AML)分析运行的夜间批量窗口及可实现恢复的结算延迟,迫使检测、解释与决策需在个位数秒的时间窗口内完成。我们提出KONTOGRAPH,这是为SEPA即时支付系统构建的端到端AML流水线,其设定的99分位决策预算为200毫秒,并报告了对1,562,860笔注入了洗钱类型且标签故意不完整的模拟支付开展的实证研究。除系统本身外,三项发现具有意义:其一,带节点级内存的时间图网络将PR-AUC较梯度提升表格基线从0.0053提升至0.1717,配对按日分块的自助法差异为+0.166,95%置信区间为[0.105, 0.241];仅节点级内存就使分数提升超一倍。其二,对每个特征仅表达一次并编译至三个执行后端,通过扰动未来的基于属性的测试强制等价性,发现了三处代码审查已通过的时点违规——每处都会夸大报告的性能。其三,对实践最重要的是,将部署的树集成导出至ONNX仅使平均分数变化了7.4×10^-8,但改变了0.26%的决策并使警报量增加了12%,因为32位累积会在3.98×10^-4的成本最优阈值附近扰动分数。我们认为,服务格式转换必须视为模型变更直至经过测量,且当候选邻域较小时,子图解释器的保真度指标可能是空泛的——我们完整报告了这一无效结果。
英文摘要
Regulation (EU) 2024/886 obliges European payment service providers to settle euro credit transfers in under ten seconds, around the clock. This removes both the overnight batch window in which anti-money-laundering (AML) analytics traditionally ran and the settlement delay that made recovery possible, forcing detection, explanation and decision inside a single-digit-second envelope. We present KONTOGRAPH, an end-to-end AML pipeline for the SEPA Instant rail built under a self-imposed 200 ms 99th-percentile budget, and report an empirical study on 1,562,860 simulated payments with injected typologies and deliberately incomplete labels. Three findings are of interest beyond the system itself. First, a temporal graph network with per-node memory improves PR-AUC over a gradient-boosted tabular baseline from 0.0053 to 0.1717, a paired day-blocked bootstrap difference of +0.166 with 95% CI [0.105, 0.241]; per-node memory alone more than doubles the score. Second, expressing each feature once and compiling it to three execution backends, with equivalence enforced by property-based tests that perturb the future, surfaced three point-in-time violations that code review had passed--each of which would have inflated reported performance. Third, and most consequential for practice, exporting the deployed tree ensemble to ONNX changed only $7.4 \times 10^{-8}$ in mean score yet altered 0.26% of decisions and inflated the alert volume by 12%, because 32-bit accumulation perturbs scores across a cost-optimal threshold of $3.98 \times 10^{-4}$. We argue that a serving-format conversion must be treated as a model change until measured, and that fidelity metrics for subgraph explainers can be vacuous when candidate neighbourhoods are small--a null result we report in full.
发表机构
- Faculty of Media Engineering and Technology(媒体工程与技术学院)
- German University in Cairo(开罗德国大学)
机构由 AI 辅助整理,请以论文原文为准。