CRISP:面向不平衡表格学习的可扩展重要性分层核心集
CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning
- University of Southern California(南加州大学)
- Coinbase
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对不平衡表格数据的大规模训练成本问题,提出线性时间核心集方法CRISP,通过重要性分层分配负类预算并加权样本,在减少93.2%训练行时保持99.7%精度,显著优于现有方法。
AI中文摘要:
大规模不平衡表格数据集使得重复的梯度提升树训练成本高昂。现有的核心集方法在移除大多数多数类样本时往往会损失精度。我们提出CRISP(通过重要性分层剪枝进行核心集约简),这是一种线性时间方法,它根据代理模型得分的分位数层分配负类预算。样本权重考虑了不相等的包含概率。在生产欺诈数据集上,当负类样本减少95%时,CRISP在约170万行(共2500万行)上进行训练,并保留了全数据平均精度的99.7%。这使总训练行数减少了93.2%。在公开的CriteoPrivateAds数据集上,在90%到99.4%的多数类缩减率下,CRISP的平均精度均值最高。在Sparkov数据集上,较低缩减率下的结果参差不齐,但在99.2%和99.4%的缩减率下,CRISP的平均精度均值最高。消融实验表明,预算分配和逆倾向加权是生产数据集增益的主要来源。
英文摘要:
Large imbalanced tabular datasets make repeated gradient-boosted tree training expensive. Existing coreset methods often lose accuracy when most majority examples are removed. We present CRISP (Coreset Reduction via Importance-Stratified Pruning), a linear-time method that allocates a negative-class budget across quantile strata of a proxy-model score. Sample weights account for unequal inclusion probabilities. At 95% negative-class reduction on a production fraud dataset, CRISP trains on approximately 1.70M of 25M rows and retains 99.7% of full-data Average Precision. This is a 93.2% reduction in total training rows. On public CriteoPrivateAds, CRISP has the highest mean Average Precision at each tested rate from 90% to 99.4% majority reduction. Sparkov results are mixed at lower rates, but CRISP has the highest mean at 99.2% and 99.4%. Ablations identify budget allocation and inverse-propensity weighting as the main sources of the production-dataset gain.