arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向薄文件档案和银行服务不足人群的可解释混合信用评分

Interpretable hybrid credit scoring for thin-file and underbanked populations

Belise Kanziga, Yaé U. Gaba, Olivier Kanamugire

arXiv 2608.26837首次发表:更新:

AI 中文总结

该研究扩展了可解释混合信用评分框架,在台湾和Zindi数据集上验证其性能,开展细粒度公平性审计,揭示社会经济维度的分配违规,可为非洲数字信贷监管提供工具。

AI 中文摘要

我们从三个方面扩展了一种残差学习混合信用评分框架(逻辑回归评分卡加上对其残差的梯度提升校正,每次预测时分解为可解释性比率ρ(x),用于衡量线性分支的贡献占比):一是基于Zindi非洲金融包容数据集(肯尼亚、卢旺达、坦桑尼亚、乌干达)的东非实证实例;二是针对框架三个可解释性区域的细粒度公平性审计;三是薄文件细分分析。在用于连续性研究的台湾信用违约基准上,经校准的混合模型达到AUC=0.776(与单独逻辑回归相比AUC提升0.057,与单独XGBoost相比提升0.001),Brier分数降低23%,且将最高违约率的借款人(69.5%)集中在完全可解释区域。在Zindi数据集上,经校准的混合模型达到AUC=0.869(与LR相比AUC提升0.015,p<0.001;与XGBoost相比降低0.004),Brier分数从0.158降至0.085(降幅46%),并复制了区域分配模式。公平性审计发现,在社会经济维度上存在严重的不透明机器学习驱动区域分配:农村受访者比城市受访者低18个百分点,初等及以下教育程度者比中等及以上教育程度者低32个百分点,乌干达受访者比肯尼亚受访者低22个百分点,而性别几乎无分配差异。该审计流程揭示了聚合公平性指标遗漏的子群体分配违规,其形式可直接被非洲各国央行的数字信贷监管机构使用。

英文摘要

We extend a residual-learning hybrid credit scoring framework (logistic regression scorecard plus a gradient-boosting correction on its residuals, decomposed at each prediction into an interpretability ratio $ρ(x)$ that measures the share attributable to the linear branch) along three axes: an East African empirical instantiation on the Zindi Financial Inclusion in Africa data (Kenya, Rwanda, Tanzania, Uganda); a fairness audit at the granularity of the framework's three interpretability regions; and a thin-file segmentation analysis. On the Taiwan Credit Default benchmark retained for continuity, the calibrated hybrid attains AUC $= 0.776$ ($Δ\mathrm{AUC} = +0.057$ vs.\ standalone logistic regression, $+0.001$ vs.\ standalone XGBoost), reduces Brier Score by 23\%, and concentrates the highest-default-rate borrowers (69.5\%) in the fully interpretable region. On Zindi, the calibrated hybrid attains AUC $= 0.869$ ($Δ\mathrm{AUC} = +0.015$ vs.\ LR, $p < 0.001$; $-0.004$ vs.\ XGBoost), cuts Brier from $0.158$ to $0.085$ (a 46\% reduction), and replicates the regional routing pattern. The fairness audit detects severe routing into the opaque ML-driven region along socioeconomic axes: rural respondents by 18 percentage points relative to urban, primary-or-less-educated by 32 points relative to secondary-and-above, and Ugandan respondents by 22 points relative to Kenyan, while gender shows essentially no routing disparity. The audit pipeline surfaces subgroup-routing violations that aggregate fairness metrics miss, in a form directly usable by African central-bank supervisors of digital credit.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑