FedIncome:数据主权约束下数字借贷中的联邦学习收入估计
FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints
浏览论文内容
中文总结 AI 辅助
针对数字借贷中收入验证缺失及数据共享受限问题,提出联邦学习框架FedIncome,在不汇集原始数据下训练共享模型,提升小样本客户收入估计精度,并改善审批决策。
中文摘要 AI 辅助
在数字贷款申请中,经过验证的收入往往不可用,这迫使贷款机构依赖申报收入,并可能导致过度放贷、过于保守的报价,或拒绝有信誉的借款人。跨机构数据共享限制使得这一问题对于训练数据有限的小型贷款机构尤为困难。我们提出了FedIncome,一个用于收入估计的联邦学习框架,使机构能够在不汇集原始借款人记录的情况下训练共享模型。使用超过一百万笔LendingClub贷款,划分为50个州级客户端,我们模拟了一个异构贷款联盟。最佳联邦模型实现了样本外R²=0.608,而集中式汇总基准为0.619。相对于集中式汇总基准,小样本客户获得的平均样本外R²提升了3.8个百分点,而拟合的客户端级关系将经验交叉点置于约4,790个训练观测值处。当汇总不可行且相关替代方案为仅本地训练时,联邦学习在所有样本量组中均提升了样本外性能,其中数据稀缺客户的收益最大。我们还将联邦收入估计与特定州和收入的债务收入比阈值相结合。在回顾性决策分析中,用联邦估计替代申报收入提高了模拟审批率,而观察到的违约率仅有适度变化。FedIncome支持数据本地化约束下的协作学习,相对于汇总训练仅有少量总体损失,相对于仅本地估计则有更大收益。
英文摘要
Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending, overly conservative offers, or rejection of creditworthy applicants. Cross-institutional data-sharing constraints make this problem especially difficult for smaller lenders with limited training data. We introduce FedIncome, a federated learning framework for income estimation that enables institutions to train a shared model without pooling raw borrower records. Using more than one million LendingClub loans partitioned into $50$ state-level clients, we simulate a heterogeneous lending consortium. The best federated model achieves out-of-time $R^2=0.608$, compared with $0.619$ for a pooled centralised benchmark. Small-sample clients obtain an average out-of-time $R^2$ improvement of $3.8$ percentage points relative to the pooled centralised benchmark, while the fitted client-level relationship places the empirical crossover at approximately $4,790$ training observations in this setting. When pooling is infeasible and the relevant alternative is local-only training, federation improves out-of-time performance across all sample-size groups, with the largest gains for data-scarce clients. We also combine federated income estimates with state- and income-specific debt-to-income thresholds. In a retrospective decision analysis, replacing reported income with the federated estimate increases simulated approval rates with only modest changes in observed default rates. FedIncome supports collaborative learning under data-locality constraints with little aggregate loss relative to pooled training and larger gains relative to local-only estimation.
发表机构
- Indian Institute of Management Indore(印多尔印度管理学院)
- Indian Statistical Institute Kolkata(加尔各答印度统计研究所)
机构由 AI 辅助整理,请以论文原文为准。