AI 中文总结
本研究构建多尺度堆叠集成预测模型并测试LLM生成信用风险解释的保真度,发现预测性能提升有限,LLM生成的解释存在事实错误,需对生成解释进行验证。
AI 中文摘要
信用评分越来越依赖于决策逻辑无法从参数中直接读取的模型,这与监管机构要求不利决策需可解释的期望存在矛盾。常见的解决方案是借助语言模型(LLM)弥合这一差距:计算特征归因,将其提供给LLM,再由LLM生成解释性理由。我们构建了这样一个端到端系统,并检验该方案后半部分的承诺是否成立。预测组件是一个多尺度堆叠集成模型,它通过在折外预测上训练的神经元学习器,将四个不同正则化的梯度提升学习器与一个残差网络进行融合。在包含32581个申请的公开信用数据集上,该模型达到测试ROC-AUC为0.9539(95%置信区间[0.9462, 0.9616]),PR-AUC为0.9137,比最优单模型的AUC提升值为0.0143(在保守独立性假设下p值为0.016)。我们的核心发现具有不对称性:排名提升是真实的,但在操作层面规模较小——在F1最优阈值下,该集成模型相比调优后的随机森林,仅在1422个案例中少漏判6个违约案例,成本加权损失降低不足2%。叙事层的失败是仅靠提示工程无法修复的:在一个经审计的案例中,模型将三个归因标记为风险增加的因素,而提供的归因却将它们标记为风险降低;模型遗漏了主导驱动因素,还引入了从未提供过的特征。我们将此归因于我们测量而非假设的属性:SHAP和LIME在重要特征上的一致性为0.80(前10个特征的重叠度),但在特征排序上的一致性较低(tau值为0.43,p值为0.18);模型最敏感输入的归因符号在不同申请者间近乎抛硬币(模态符号占比为0.53);校准度(ECS=0.117)和扰动稳定性(DPD=0.078)均未达到我们设定的阈值。受限提示是必要的,但并非充分条件:生成后的依据必须经过验证,而非假设其正确。
英文摘要
Credit scoring increasingly relies on models whose decision logic cannot be read off their parameters, in tension with supervisory expectations that adverse decisions be explainable. A common proposal closes that gap with a language model: compute feature attributions, hand them to an LLM, and let it write the rationale. We build such a system end to end and test whether the second half of the promise holds. The predictive component is a multi-scale stacking ensemble fusing four differently regularised gradient-boosting learners with a residual network through a neural meta-learner trained on out-of-fold predictions. On a public 32,581-application credit dataset it reaches test ROC-AUC 0.9539 (95% CI [0.9462, 0.9616]) and PR-AUC 0.9137, beating the best single model by Delta-AUC = 0.0143 (p = 0.016 under a conservative independence assumption). Our central finding is asymmetric. The ranking gain is real but operationally small: at the F1-optimal threshold the ensemble avoids only six additional missed defaults out of 1,422 against a tuned random forest, cutting cost-weighted loss by under 2%. The narrative layer fails in a way prompt engineering alone does not fix. In an audited case the model named three factors as risk-increasing that the supplied attributions scored as risk-reducing, omitted the dominant driver, and introduced a feature never given to it. We trace this to properties we measure rather than assume: SHAP and LIME agree on which features matter (overlap@10 = 0.80) but not on their order (tau = 0.43, p = 0.18), and the attribution sign for the model's most sensitive input is near a coin flip across applicants (modal-sign share 0.53). Calibration (ECS = 0.117) and perturbation stability (DPD = 0.078) both fall short of our own thresholds. Constrained prompting is necessary but not sufficient: grounding must be verified after generation, not assumed.
Comments13 pages, 9 figures, 5 tables