arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分类器何时能助力大语言模型?用于信用卡违约预测的分类器引导提示及混合分类器-大语言模型模型

When Does a Classifier Help an LLM? Classifier-Guided Prompting and Hybrid Classifier-LLM Models for Credit-Default Prediction

Rishi Datta, Lavanya Prahallad

arXiv 2608.30086首次发表:更新:

发表机构

Amador Valley High School; Research Spark Hub Inc.(阿马多尔谷高中; Research Spark Hub 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对信用卡违约预测任务,对比了单模型性能,提出将分类器的预测概率加入LLM提示的方法,可提升其AUC-ROC至与随机森林相当,同时保持更高召回率,推荐使用该分类器引导提示。

AI 中文摘要

信用卡违约预测是金融决策中的一项重要任务。传统方法会在表格型特征上使用拟合后的分类器,如逻辑回归和随机森林。近来,大语言模型(LLM)通过提示被应用于该任务。本研究探讨如何将拟合后的分类器与LLM结合用于信用卡违约预测,区分了指示LLM模仿分类器与利用分类器构建提示这两种方式,假设拟合后的分类器可提供LLM提示所欠缺的排序能力。我们在信用卡客户违约数据集(Default of Credit Card Clients dataset)上开展实验,报告了召回率、F1值,以及ROC曲线和精确率-召回率曲线下面积,并给出了自助法置信区间。结果显示,少样本LLM在所有单模型中具有最高的召回率(0.47)和F1值(0.50),但AUC-ROC(0.72)低于随机森林(0.79);指示LLM模仿分类器未产生显著变化;将提示剪枝为分类器的8个最重要特征,可使召回率提升0.071、F1值提升0.032;将分类器的预测概率加入提示,可使LLM的AUC-ROC从0.72提升至0.78,与随机森林相当,同时保持比随机森林高0.118的召回率;反向组合方式及使用多个分类器均无帮助,因此我们推荐采用简单的分类器引导提示用于基于LLM的信用预测。

英文摘要

Credit-default prediction is an important task in financial decision making. Traditional methods use fitted classifiers such as logistic regression and random forests on tabular features. Large language models (LLMs) have recently been applied to this task through prompting. In this work we study how a fitted classifier and an LLM can be combined for credit-default prediction. We distinguish telling the LLM to imitate a classifier from using the classifier to build the prompt. We hypothesize that a fitted classifier can supply the ranking ability that an LLM prompt lacks. We experiment on the Default of Credit Card Clients dataset, and report recall, F1, and the area under the ROC and precision-recall curves, with bootstrap confidence intervals. We observe that a few-shot LLM has the highest recall (0.47) and F1 (0.50) of any single model but ranks worse than a random forest (AUC-ROC 0.72 against 0.79). Instructing the LLM to imitate a classifier gives no significant change. Pruning the prompt to the classifier's eight most important features raises recall by 0.071 and F1 by 0.032. Adding the classifier's predicted probability to the prompt raises the LLM's AUC-ROC from 0.72 to 0.78, matching the random forest, while keeping 0.118 higher recall than it. The reverse composition, and the use of several classifiers, do not help. We thus recommend a simple classifier-guided prompt for LLM-based credit prediction.

Comments6 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑