FraudBench:面向金融风险评估的协议敏感型对抗鲁棒性基准测试
FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment
AI总结:
FraudBench是面向金融风险评估的协议敏感型对抗鲁棒性基准,通过三种匹配协议评估不同模型,发现鲁棒性结论具协议敏感性,提出应将领域约束纳入攻击生成。
AI中文摘要:
机器学习模型被广泛应用于金融欺诈与信用风险检测,但其对抗鲁棒性仍难以评估,因为金融表格数据涉及领域特定约束、严重的类别不平衡以及攻击者能力的不对称性。本文认为,在该场景下,鲁棒性不仅是模型的属性,也是评估协议的属性;不同的约束实施方式和攻击者能力会导致截然不同的鲁棒性结论。本文提出了FraudBench——一种用于金融欺诈与信用风险检测的协议敏感型对抗鲁棒性基准测试。FraudBench未将领域约束视为事后有效性检查,而是在三种匹配协议下评估相同的数据集-模型-攻击-防御设置:无约束攻击、事后可行性过滤以及部署感知的约束集成攻击。FraudBench涵盖四个公开金融数据集,并使用三种攻击设置评估神经、基于树的以及集成模型。实验结果表明,鲁棒性结论具有高度的协议敏感性:在白盒设置下的Lending Club贷款数据中,事后过滤平均仅保留3.7个可行翻转样本,而在相同扰动预算下,带有攻击者可变性掩码的攻击内投影则产生2832.3个可行翻转样本;针对IEEE-CIS的结果进一步表明,可行性与攻击者能力是相互独立的维度,而黑盒评估则显示协议选择可改变模型家族的排名。这些发现提示,欺诈鲁棒性评估应同时报告预测性能下降与攻击可行性,并应将领域约束纳入攻击生成过程,而非将其作为事后处理检查。
英文摘要:
Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate because financial tabular data involve domain-specific constraints, severe class imbalance, and asymmetric attacker capability. We argue that, in this setting, robustness is not only an attribute of the model, but also an attribute of the evaluation protocol. Different ways of enforcing constraints and capability can lead to substantially different robustness conclusions. This paper presents FraudBench, a protocol-sensitive benchmark for adversarial robustness evaluation in financial fraud and credit-risk detection. Rather than treating domain constraints as post-hoc validity checks, FraudBench evaluates the same dataset--model--attack--defence setting under three matched protocols: unconstrained attacks, post-hoc feasibility filtering, and deployment-aware constraint-integrated attacks. FraudBench covers four public financial datasets, and evaluates neural, tree-based, and ensemble models using three attack settings. Our results show that robustness conclusions are highly protocol-sensitive. On Lending Club Loan Data under the white-box setting, post-hoc filtering leaves only 3.7 feasible-flipped examples on average, whereas in-attack projection with attacker mutability masking produces 2,832.3 feasible-flipped examples under the same perturbation budget. The results on IEEE-CIS further show that feasibility and attacker capability are separate axes, while black-box evaluation shows that protocol choice can alter model-family rankings. These findings suggest that fraud robustness evaluation should report predictive degradation and attack feasibility jointly, and should incorporate domain constraints into attack generation rather than treating them as post-processing checks.