arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

信用卡欺诈检测中的类别加权与金额条件设定:基于美元度量的时间解释性审计研究

Class Weighting versus Amount Conditioning in Credit-Card Fraud Detection: A Dollar-Metric Study with a Temporal Explanation Audit

Chenyu Wu

arXiv 2607.14686首次发表:更新:

AI 中文总结

研究信用卡欺诈检测中交易金额对训练权重及警报排序的作用,通过固定总欺诈案例权重改变分配方式,用XGBoost在两个数据集测试多种方法,结果表明金额衍生特征重要,金额作加权规则效果不佳,作特征和排序变量有用。

AI 中文摘要

信用卡欺诈损失以货币计算,但论文常以交易级分数评判模型。我们探讨交易金额应塑造训练权重还是用于后续警报排序。为将此问题与普通类别不平衡处理区分开,我们保持总欺诈案例权重固定,仅改变其在欺诈案例间的分配。实验在两个按时间顺序排列的信用卡欺诈数据集上,用XGBoost测试了无加权训练、标准类别加权、匹配对数金额加权、更强的金额加权变体以及分数乘以金额重新排序等方法。指标包括平均精度、美元召回率和固定警报预算下的美元精度,在五个种子上进行测试,并对主要对比采用95%的日块自举区间。结果比预期更窄。金额衍生的比率和速度特征携带了大部分信号,而一旦这些特征在模型中,原始金额字段增加的信息很少。在匹配设置中,金额条件训练比类别加权仅带来小的增益,且不能始终击败普通无加权模型。更强的金额权重能追回更多欺诈美元,但排序质量和美元精度较低。训练后按分数乘以金额重新排序警报带来最大的美元召回率变化。一项小型SHAP审计发现,欺诈案例的逐月归因变动比总体流量更大。在这些测试中,金额作为特征和警报排序变量是有用的,但本身并非更好的样本加权规则。

英文摘要

Credit-card fraud losses are monetary, but papers often judge models with transaction-level scores. We ask whether transaction amount should shape training weights or be used later to order alerts. To separate this question from ordinary class imbalance handling, we keep total fraud-case weight fixed and vary only its allocation across fraud cases. The experiments test two chronological card-fraud datasets with XGBoost under unweighted training, standard class weighting, matched log-amount weighting, stronger amount-weighted variants, and score times amount reranking. Metrics are average precision, dollar recall, and dollar precision at fixed alert budgets over five seeds, with 95 percent day-block bootstrap intervals for the main contrasts. Results are narrower than expected. Amount-derived ratio and velocity features carry much of the signal, while raw amount fields add little once those features are in the model. In the matched setting, amount-conditioned training gives only small gains over class weighting and does not consistently beat the plain unweighted model. Stronger amount weights recover more fraudulent dollars, but at lower ranking quality and dollar precision. Reranking alerts by score times amount after training gives the largest dollar-recall shift. A small SHAP audit finds larger month-to-month attribution movement for fraud cases than for aggregate traffic. In these tests, amount is useful as a feature and as an alert-ordering variable, not by itself as a better sample-weighting rule.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑