发表机构
Queensland University of Technology(昆士兰科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过45项不平衡比例各异的二分类任务发现,基于单个信用卡欺诈数据集得出的阈值调优与重采样结论不可靠,验证集校准误差无法预测调优增益,公开了相关工具链与指标。
AI 中文摘要
不平衡数据处理通常在单个基准数据集上进行评估,所得结论被当作该方法的固有属性,我们证明这种做法是不可靠的。在公开的Kaggle信用卡欺诈数据集上,采用无泄露嵌套交叉验证协议(在保留的内部验证折上选择决策阈值),普通随机森林在默认0.5阈值下的F1值为0.861±0.021,而阈值调优未给其带来增益(F1差值为-0.002)。仅据此会得出一个颇具吸引力的结论:对于校准良好的集成模型,不平衡处理是不必要的。随后,我们将相同协议应用于45项二分类任务,涵盖不平衡比例从1:1.5到1:178(共2025次模型拟合,涉及4个模型族),结论发生反转。随机森林在整套任务中从阈值调优中获益最多(F1差值为+0.101±0.134),同时其他三个模型族几乎完全复现了其在欺诈数据集上的表现。SMOTE同样对欺诈数据集有害,但在整套任务中表现出增益(平均F1差值为+0.076;138次获胜,39次失败;Wilcoxon符号秩检验p值为2.7e-17)。另有两项结果:阈值调优的增益与不平衡比例呈非单调关系:1:5以下接近零,在1:15-1:40区间达到峰值+0.120,在1:100以上降至+0.045——这解释了为何不平衡比例为1:577的欺诈数据集并非研究该问题的代表性场景。此外,我们否定了一种直观启发式方法:验证集校准误差无法预测调优增益(预期校准误差相关系数r=-0.087;Brier分数相关系数r=+0.137),因此校准诊断无法告知从业者是否值得进行阈值调优。我们公开了该协议、45项任务的工具链及所有单次运行的指标。
英文摘要
Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method. We show this practice is unsafe. On the public Kaggle credit-card fraud dataset, under a leakage-free nested cross-validation protocol in which the decision threshold is selected on a held-out inner validation fold, a plain Random Forest at the default 0.5 threshold attains F1 = 0.861 +/- 0.021, and threshold tuning yields it no benefit (delta-F1 = -0.002). Read alone, this supports an appealing conclusion: for a well-calibrated ensemble, imbalance handling is unnecessary. We then apply the identical protocol to 45 binary tasks spanning imbalance ratios from 1:1.5 to 1:178 (2,025 model fits, four model families). The conclusion reverses. Random Forest benefits most from threshold tuning across the suite (delta-F1 = +0.101 +/- 0.134), not least, while three other families replicate their fraud-dataset behaviour almost exactly. SMOTE likewise harms the fraud dataset but helps across the suite (mean delta-F1 = +0.076; 138 wins, 39 losses; Wilcoxon p = 2.7e-17). Two further results. Threshold-tuning benefit is non-monotonic in the imbalance ratio: near zero below 1:5, peaking at +0.120 in the 1:15-1:40 band, declining to +0.045 beyond 1:100 - explaining why the fraud dataset, at 1:577, is an unrepresentative place to study the question. And we reject an intuitive heuristic: validation-set calibration error does not predict tuning benefit (expected calibration error r = -0.087; Brier r = +0.137), so calibration diagnostics cannot tell a practitioner whether tuning is worthwhile. We release the protocol, the 45-task harness, and all per-run metrics.
Comments13 pages, 3 figures, 5 tables. Code and per-run metrics released