AI 中文总结
该研究针对自动信贷违约预测,发现高收入违约者被误判为噪声的比例更高,通过序列特征盲化分离出收入、利率等三类差异驱动因素,揭示盲化敏感属性无法保证公平性,为金融AI审计提供依据。
AI 中文摘要
以数据为中心的整理流程常依赖模型置信度分数标记和过滤有噪声或标签错误的训练样本。在大规模消费者借贷样本(LendingClub,样本量N=1,344,936)上评估该过滤惯例时,发现了潜在的人口统计学不对称性:高收入违约者被归类为标签噪声的比例远高于低收入违约者(克莱默V值约为0.03-0.07)。从平等机会[Hardt等人,2016]的视角重新审视这一行为,会发现更严重的差异:最终违约的高收入与低收入借款人之间的真正例率(召回率)存在16.86个百分点的差距。采用序列特征盲化方法,我们得以在三种不同机制中分离出该差异的驱动因素:(1)直接依赖申请人自报收入;(2)算法吸收了初始利率中编码的上游机构偏见;(3)残留差异(交叉验证中为3.55个百分点,保留测试分区中为2.56个百分点,Z=-4.04,p<0.0001),即使模型剔除收入和利率后仍存在。样本外符号SHAP估值表明,该残留差距由结构代理变量维持,最显著的是贷款金额和住房拥有状态。这些实证结果表明,当机构定价决策和行为代理变量共同重构被遗漏信号时,仅让算法对敏感属性盲化无法确保公平性。我们概述了这些发现对受监管金融机构内以数据为中心的AI工作流审计的实际意义。
英文摘要
Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances. Evaluating this filtering convention on a large-scale consumer lending sample (LendingClub, N = 1,344,936) uncovers an underlying demographic asymmetry: high-income defaulters are disproportionately classified as label noise relative to low-income defaulters (Cramer's V approximately 0.03-0.07). Re-examining this behavior through the lens of equal opportunity [Hardt et al., 2016] reveals a far more severe discrepancy: a 16.86 percentage point gap in true positive rate (recall) between high- and low-income borrowers who ultimately defaulted. Implementing a sequential feature-blinding methodology allows us to isolate the drivers of this disparity across three distinct mechanisms: (1) direct reliance on self-reported applicant income; (2) algorithmic absorption of upstream institutional bias encoded within origination interest rates; and (3) a residual disparity (3.55 percentage points in cross-validation; 2.56 percentage points on a held-out test partition, Z = -4.04, p < 0.0001) that remains even after purging both income and interest rates from the model. Out-of-sample signed SHAP valuations demonstrate that this residual gap is maintained by structural proxies, most notably loan amount and home ownership status. These empirical findings show that simply blinding an algorithm to sensitive attributes fails to ensure fairness when institutional pricing decisions and behavioral proxy variables collectively reconstruct the omitted signals. We outline the practical implications of these findings for auditing data-centric AI workflows within regulated financial institutions.
Comments8 pages, 2 figures