发表机构
Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究欺诈检测系统,通过分层管道结合多种方法,包括梯度提升分类器等。在PaySim数据集上实验,各组件在特定条件下起作用,如校正后图特征等未提升全测试集平均精度,但在部分子集有作用,调查代理表现不佳,还标记了其错误。
AI 中文摘要
欺诈检测系统必须随着交易量的增加而扩展,同时保持可解释性和可审查性。我们研究了PaySim数据集上的分层管道,该管道结合了梯度提升分类器、图衍生的结构特征、基于自动编码器的异常信号、TreeSHAP解释以及应用于分类器评分不确定的案例的有界语言模型调查代理。在进行任何模型比较之前,我们识别并消除了模拟器特定的平衡捷径,否则会夸大基线性能。校正后,图特征和异常信号均未提高完整测试集上的平均精度。然而,在获得中间基线分数的案例子集中,两者对欺诈的排名都更好。在注入多账户欺诈环的对照实验中,工程结构特征恢复了所有注入的测试交易,而表格基线遗漏了大约四分之一。调查代理的表现不如其依赖的分类器的直接阈值,在平衡的60个案例样本中,准确率为65.0%,而直接阈值为71.7%,尽管可以访问模型解释、图上下文和检索到的参考案例。在代理更改的八个决策中,六个用错误替换了正确的分类器输出,并且在每种情况下都产生了连贯的书面理由。基于探索性分歧的升级规则标记了其中两个代理错误以供人工审查,而未标记任何正确决策。我们得出结论,分层欺诈系统的每个组件仅在特定条件下才起作用,并且调查代理看似合理的理由并不证明决策更好。
英文摘要
Fraud detection systems must scale with rising transaction volume while remaining explainable and reviewable. We study a layered pipeline on the PaySim dataset that combines a gradient-boosted classifier, graph-derived structural features, an autoencoder-based anomaly signal, TreeSHAP explanations, and a bounded LLM investigation agent applied to cases the classifier scores uncertainly. Before any model comparison, we identify and remove a simulator-specific balance shortcut that would otherwise inflate baseline performance. After this correction, neither the graph features nor the anomaly signal improves Average Precision on the full test set. Both, however, rank fraud better within the subset of cases receiving intermediate baseline scores. In a controlled experiment with injected multi-account fraud rings, engineered structural features recover all injected test transactions, while the tabular baseline misses roughly a quarter of them. The investigation agent underperforms direct thresholding of the classifier it relies on, reaching 65.0% accuracy against 71.7% on a balanced 60-case sample, despite having access to model explanations, graph context, and retrieved reference cases. Of the eight decisions the agent changed, six replaced correct classifier outputs with errors, and it produced a coherent written rationale in each case. An exploratory disagreement-based escalation rule flagged two of these agent errors for human review without flagging any correct decision. We conclude that each component of a layered fraud system contributes only under specific conditions, and that a plausible rationale from an investigation agent is not evidence of a better decision.
Comments14 pages, 3 figures. Code available at https://github.com/rahil1303/auditable-fraud-investigation