当大型语言模型的语言置信度与内部置信度出现偏差时
When Linguistic and Internal Confidence Diverge in Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究对比大型语言模型的语言置信度与内部置信度,发现二者常出现偏差,提出需用多维度诊断评估语言置信度后再用于下游可靠性流程。
中文摘要 AI 辅助
用户常要求大型语言模型(LLMs)报告自身的置信度,但目前尚不清楚这类语言置信度是否与模型的内部置信度相符。本研究针对8个分类任务、2个生成任务以及3个系列的30个模型展开该问题的探究。对于分类任务,从关联度、量级一致性、校准度三个维度将语言置信度与基于logits的置信度进行对比;对于生成任务,测试语言置信度是否与基于语义熵的不确定性相符。结果显示,二者在上述维度频繁出现偏差:实例级关联度平均水平较弱,不过在较简单的任务项以及性能更强的基础模型中会有所提升;经过指令微调的模型通常报告更高的置信度,有时关联度也更高,但同时存在更大的置信度差距与更差的校准度;提示工程主要改变报告置信度的分布,态度线索会夸大置信度却未提升一致性,而分数范例在避免置信度值坍塌时可保留排序信号。回归分析表明,置信度分数的分布属性可解释大部分观测到的一致性模式,在控制变量后模型元数据的影响较小。这些结果支持语言置信度的有损信道观点:更分散的 verbal 置信度分布可携带有用的排序信息,但无法使分数得到校准。因此,语言置信度在用于下游可靠性流程前,应通过多维度诊断进行评估。
英文摘要
Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 generation tasks and 30 models from three families. For classification, we compare linguistic confidence with logits-based confidence along three axes: association, magnitude agreement and calibration. For generation, we test whether linguistic confidence tracks semantic-entropy-based uncertainty. The axes frequently diverge. Instance-level association is weak on average, although it improves on easier items and for stronger base models. Instruction-tuned models often report higher confidence and sometimes show higher association, but they also have larger confidence gaps and worse calibration. Prompt design mostly changes the distribution of reported confidence. Attitude cues inflate confidence without improving alignment, while score exemplars can preserve rank-order signal when they avoid collapsed confidence values. Regression analyses show that distributional properties of confidence scores explain much of the observed alignment pattern, with model metadata playing a smaller role after controls. These results support a lossy-channel view of linguistic confidence. A more dispersed verbal confidence distribution can carry useful rank information, but it does not make the scores calibrated. Linguistic confidence should therefore be evaluated with multi-axis diagnostics before being used in downstream reliability pipelines.
发表机构
- Dartmouth College(达特茅斯学院)
- Oakland University(奥克兰大学)
机构由 AI 辅助整理,请以论文原文为准。