利用大语言模型传达信用风险:基于标准数据与替代数据模型的解释评估
Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models
- Western University(西安大略大学)
- University of Southampton(南安普顿大学)
- Reykjavik University(雷克雅未克大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究探究大语言模型能否作为解释层转化信用风险模型的事后解释,构建三类流程并采用三类LLM,发现证据表征是解释质量的关键制约,专业人士对证据标准要求更严格。
AI中文摘要:
信用决策是一项高风险任务,模型输出必须准确且可解释,以支持合规决策。尽管极端梯度提升(XGBoost)和图神经网络(GNN)等现代信用风险模型提升了预测性能,但它们的解释通常对利益相关者而言过于技术化,造成了可能影响审批、拒贷及公平性判断的沟通缺口。本研究探讨大语言模型(LLM)是否可作为解释层,将事后解释产物转化为利益相关者适用的风险叙事。利用房地美(Freddie Mac)的单户贷款层面数据,我们构建了三条流程:标准表格型流程(XGBoost + SHAP)、两条基于替代数据的流程,分别为纯网络型流程(GNN + GNNExplainer)和双模态流程(结合表格与网络数据)。我们采用三种LLM配置生成叙事:小型微调LLM(Gemma 3 4B)、大型微调LLM(DeepSeek R1 70B)以及零样本商业LLM(Gemini 2.5)。通过对所有流程的自动检查,以及针对双模态解释的人类研究,从八个决策相关维度对比信用风险专业人士与非专业人士,对解释质量进行评估。我们得出三项主要发现:第一,流程对证据基础分数方差的解释高于语言模型,即解释质量的关键制约因素是证据表征,而非所用模型;第二,解释叙事能可靠地指出影响因素,但在说明影响方向时可靠性较低,这可能对不利行动沟通产生影响;第三,专业人士比非专业人士适用更严格的证据标准。我们讨论了风险模型治理的启示,包括部署考量以及领域适配LLM在受监管信用场景中的价值。
英文摘要:
Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.