PaMIR:公共信用违约数据集的开放基准
PaMIR: Open Benchmark of Public Credit-Default Datasets
AI总结:
PaMIR是一个开放基准,汇集19个公共信用违约数据集(124万样本),在标签稀缺且延迟到达场景下,通过到达时评分和标签预算报告AUC,为信用违约预测提供标准化评估。
AI中文摘要:
我们发布了PaMIR(用于风险推理的公共按到达顺序测量),这是一个在标签稀缺且延迟到达时用于信用违约预测的开放基准。该领域的参考基准研究各使用八个数据集,其中只有两个或四个是公开的。PaMIR汇集了19个带有二元违约标签的公共数据集——涵盖来自九个国家的124万笔贷款、企业和信用卡账户——这些数据由固定的源快照通过一个经泄漏审计的流程重建,且从未重新分发;据我们所知,截至今日,它是同类中唯一的基准。每个模型都是一个单一函数,在重复的独立同分布划分和标签延迟流下进行评分,其中每个申请在到达时即被评分,并按标签预算报告AUC;除非每个数据集都被评分,否则不公布整体均值。一个合成数据测试工具在不允许生成器看到保留行的情况下测试生成的训练行。本报告描述了这一动态基准的0.4.0版本。
英文摘要:
We release PaMIR (Public Arrival-ordered Measurement for Inference in Risk), an open benchmark for credit-default prediction when labels are scarce and arrive late. The field's reference benchmark studies use eight datasets each, only two or four of them public. PaMIR brings together 19 public datasets with binary default labels -- 1.24M loans, firms and card accounts from nine countries -- rebuilt from pinned source snapshots by one leakage-audited recipe and never redistributed; to our knowledge it is the one of its kind as of today. Every model is a single function, scored under a repeated i.i.d. split and a label-delayed stream in which each application is scored on arrival, with AUC reported by label budget; fleet means are withheld unless every dataset is scored. A synthetic-data harness tests generated training rows without letting a generator see held-out rows. This report describes release 0.4.0 of this living benchmark.