大语言模型中置信度的计算基础
The Computational Basis of Confidence in Large Language Models
浏览论文内容
中文总结 AI 辅助
研究大语言模型中置信度的计算基础,利用统计决策置信度框架,通过答案 - 对数差异测试其预测特征,结果表明在多种任务中答案对数可作潜在决策变量读出,为多模态语言模型置信度提供解释并统一研究框架。
中文摘要 AI 辅助
可靠的置信度(即模型自身答案正确的概率)对于语言模型的可靠部署至关重要。现有工作大多通过预测正确性的程度和是否校准来评估置信度,而一个更根本的问题仍然存在:置信度信号本身代表什么?答案对数可能反映一个足以计算规范置信度的潜在决策变量,或者是一个以非贝叶斯方式组合现有证据的启发式偏好信号。我们使用统计决策置信度(SDC,一种来自计算神经科学的规范框架)来解决这个问题。将答案 - 对数差异(LD)视为潜在决策变量的候选读出,我们测试了SDC预测的定性特征。在三个感知辨别任务和一个基于记忆的决策任务中,跨越三个多模态非推理模型和一个推理模型,LD满足这些特征,包括诊断正确/错误折叠 - X模式,表明在这些设置中,答案对数表现为潜在决策变量的单调读出,而不是启发式偏好分数。在复杂视觉推理中,LD继续预测正确性超过客观任务难度,但SDC的完整几何特征不存在,说明了当明确的规范过程模型不可用时该框架的当前边界。这些结果提供了对多模态语言模型中置信度的计算解释,描绘了答案对数何时表现为潜在决策变量的读出,并将SDC确立为研究生物和人工智能中置信度的统一框架。
英文摘要
Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is calibrated, leaving open a more fundamental question: what does the confidence signal itself represent? Answer logits may reflect a latent decision variable sufficient to compute normative confidence, or instead a heuristic preference signal that combines the available evidence in a non-Bayesian manner. We address this using statistical decision confidence (SDC), a normative framework from computational neuroscience. Treating the answer-logit difference (LD) as a candidate readout of the latent decision variable, we test the qualitative signatures predicted by SDC. Across three perceptual discrimination tasks and a memory-based decision task, spanning three multimodal non-reasoning models and one reasoning model, LD satisfied these signatures -- including the diagnostic correct/error folded-X pattern -- showing that, in these settings, answer logits behave as monotonic readouts of a latent decision variable rather than heuristic preference scores. In complex visual reasoning, LD continued to predict correctness beyond objective task difficulty, but the full geometric signatures of SDC were absent, illustrating the current boundary of the framework when explicit normative process models are unavailable. These results provide a computational account of confidence in multimodal language models, delineate when answer logits behave as readouts of a latent decision variable, and establish SDC as a unifying framework for studying confidence across biological and artificial intelligence.
发表机构
- Google DeepMind(谷歌DeepMind)
- École Polytechnique(巴黎综合理工学院)
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。