自信地犯错:从大语言模型内部状态中检测金融问答中的幻觉
Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States
浏览论文内容
中文总结 AI 辅助
研究金融应用中LLMs自信却错误答案(自信幻觉)的检测,通过在残差流训练线性探测器并在FinQA和TAT-QA基准评估,发现探测器检测幻觉优势明显,可作为经济高效分类机制用于高风险金融应用答案审查。
中文摘要 AI 辅助
在金融应用中,大语言模型(LLMs)在给出自信却错误的答案时后果最为严重。有保留、不确定的答案会引发审查,而自信的错误会在毫无预警的情况下悄悄影响下游决策。我们研究能否从模型的内部激活中可靠地检测出这些自信的错误答案(即自信幻觉),以及这些激活是否携带了超出其可观察输出的信息。我们在残差流上训练线性探测器,并在两个基于真实文件构建的问答基准(FinQA和TAT-QA)上进行评估。行为置信度通过对同一问题的八个重采样答案之间的一致性来衡量,并将探测器的有效性与诸如令牌对数概率和模型对其答案的真假自我评估等基线进行比较。我们的研究结果表明,在自信答案中,八个重采样答案都一致的那些答案中,在FinQA上有15%-23%是错误的。在检测幻觉方面,探测器相对于基线方法具有显著优势,在Qwen3-8B、Llama-3.1-8B和Gemma-2-9B模型上,探测器的AUROC为0.68-0.77,而最佳基线则降至0.55-0.63。我们的结果表明,探测可以成为一种经济高效的分类机制,用于在高风险金融应用中将大语言模型的答案路由到人工审查和质量控制程序中。
英文摘要
Large language models (LLMs) in financial applications fail most consequentially when they are confidently wrong. Hedged, uncertain answers invite scrutiny, whereas confident errors silently degrade downstream decisions without warning. We ask how reliably such confidently wrong answers, or confident hallucinations, can be detected from a model's internal activations, and whether those activations carry information beyond its observable outputs. We train linear probes on the residual stream and evaluate them on two established question-answering (QA) benchmarks built from real filings, FinQA and TAT-QA. Behavioral confidence is measured as the agreement among eight resampled answers to the same question, and probe effectiveness is compared against baselines, such as token log-probabilities and the model's own True/False self-assessment of its answer. Our findings show that among confident answers, those for which all eight resamples agree, 15-23% are wrong on FinQA. There the probes have a significant advantage over baseline methods in detecting hallucinations, holding 0.68-0.77 AUROC while the best baselines fall to 0.55-0.63, across Qwen3-8B, Llama-3.1-8B, and Gemma-2-9B. Our results suggest that probing can be a cost-effective triage mechanism for routing LLM answers to human review and quality control procedures in high-stakes financial applications.
发表机构
- St. John Fisher University(圣约翰费舍尔大学)
机构由 AI 辅助整理,请以论文原文为准。