发表机构
Ant International(蚂蚁国际)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对金融机器学习中信用评分冷启动问题,提出模型无关的排序先验对齐框架,在多数据集及多模型上验证其有效性,且对齐增益随数据、模型容量及特征质量降低而提升。
AI 中文摘要
冷启动信用评分——即部署带有稀缺标注数据、弱特征或最小容量的模型——是金融机器学习中反复出现的问题。当新的贷款产品推出时,标注的违约数据稀缺,特征管道不成熟,且模型必须以最小容量部署以避免过拟合。标准防御措施在相同的有限数据上运行;所需的是基于领域知识的外部正则化来源。我们提出排序先验对齐(Ranking Prior Alignment),这是一种模型无关框架,通过温度缩放的KL散度损失将来自领域专家、教师模型或大语言模型(LLM)的外部排序先验提炼到任何评分模型中。该框架通过单一公式统一了神经(多实例学习注意力,MIL attention)和基于树的(XGBoost自定义目标)架构:L = L_task + γ(t) * KL(P_agent || P_model),其中γ(t)遵循指数衰减计划。该方法在推理时不需要外部模型,其基于树的实例可容忍高达η=0.5的标注噪声。在超过150万商家的工业数据集上,MIL对齐在3000个包(1个分布内ID + 3个分布外OOT + 3个退化指标;在OOT-1上的峰值AUC差值为+0.020)实现了7/7个正评估单元,XGBoost消融在300个包时实现了9/9个正指标。在公开的Amex数据集上的跨数据集验证显示5/5个正折(平均AUC差值为+0.041)。四个模型家族(MIL、XGBoost、LightGBM、逻辑回归)和四个教师架构表明,该框架既是模型无关的,也是先验源无关的。我们进一步观察到,对齐增益呈现反缩放模式:随着数据丰度N、模型容量C和特征质量Q的降低,收益会增加,这有助于从业者决定何时投入先验标注。
英文摘要
Cold-start credit scoring -- deploying models with scarce labeled data, weak features, or minimal capacity -- is a recurring problem in financial machine learning. When a new lending product launches, labeled default data is scarce, feature pipelines are immature, and models must be deployed with minimal capacity to avoid overfitting. Standard defenses operate on the same limited data; what is needed is a source of external regularization grounded in domain knowledge. We propose Ranking Prior Alignment, a model-agnostic framework that distills external ranking priors (from domain experts, teacher models, or LLMs) into any scoring model via a temperature-scaled KL divergence loss. The framework unifies neural (MIL attention) and tree-based (XGBoost custom objective) architectures through a single formulation: L = L_task + gamma(t) * KL(P_agent || P_model), where gamma(t) follows an exponential decay schedule. The method requires no external model at inference, and its tree-based instantiation tolerates annotation noise up to eta = 0.5. On an industrial dataset of over 1.5M merchants, MIL alignment achieves 7/7 positive evaluation cells at 3K bags (1 ID + 3 OOT + 3 degradation metrics; peak Delta AUC = +0.020 on OOT-1), and XGBoost ablation achieves 9/9 positive metrics at 300 bags. Cross-dataset validation on public Amex shows 5/5 positive folds (avg Delta AUC = +0.041). Four model families (MIL, XGBoost, LightGBM, Logistic Regression) and four teacher architectures show that the framework is both model-agnostic and prior-source-independent. We further observe that alignment gains exhibit an inverse-scaling pattern: benefits grow as data abundance N, model capacity C, and feature quality Q decrease, helping practitioners decide when to invest in prior annotation.
Comments8 pages, 5 figures, 14 tables