发表机构
SiftyML, LTD(SiftyML有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较自回归语言模型的局部与全局置信度,发现二者相关性弱且与正确性关联不同,分歧可诊断采样不稳定性,故置信度应视为明确定义的测量而非内在属性。
AI 中文摘要
现代预测系统会暴露多种通常被解释为置信度度量的量。然而,这些量可能概括了预测过程的不同方面。当置信度被用于评估可靠性或为下游监督和控制提供信息时,这种区别至关重要。我们通过比较局部置信度(定义为贪心选择答案标记的概率)与全局置信度(定义为重复采样下众数答案的频率),研究在自回归语言模型中不同的置信度读数是否在经验上可互换。在MMLU和ARC Challenge上,这两种信号相关性较弱,并且在与正确性的关联上存在显著差异:全局置信度与正确性中等相关,而局部置信度几乎不相关。我们进一步测试了问题级别的信号间分歧是否与采样不稳定性相关。在ARC上,较大的局部-全局置信度差距与更高的答案熵、更多不同的采样答案以及更低的众数答案集中度相关。当分歧和不稳定性由不相交的随机样本估计时,差距-熵关联仍然存在,表明它并非由共享的有限样本变异所解释。在MMLU上,相应的关系明显较弱,只有4%的问题表现出采样不稳定性。这些结果表明,从同一预测系统导出的置信度读数在经验上不可互换,并且它们的分歧可以作为不稳定采样行为的诊断指标。因此,置信度应被视为一种明确定义的测量,而非模型单一的内在标量属性,尤其是在用于下游评估、监督或控制时。
英文摘要
Modern predictive systems expose multiple quantities that are commonly interpreted as measures of confidence. However, these quantities can summarize different aspects of the predictive process. This distinction matters when confidence is used to evaluate reliability or inform downstream oversight and control. We investigate whether different confidence readouts are empirically interchangeable in an autoregressive language model by comparing local confidence, defined from the probability of the greedy-selected answer token, with global confidence, defined from modal-answer frequency under repeated sampling. Across MMLU and ARC Challenge, the two signals are weakly correlated and differ substantially in their association with correctness: global confidence is moderately associated with correctness, whereas local confidence shows little association. We further test whether question-level disagreement between the signals is associated with sampling instability. On ARC, larger local--global confidence gaps are associated with higher answer entropy, more distinct sampled answers, and lower modal-answer concentration. The gap--entropy association persists when disagreement and instability are estimated from disjoint stochastic samples, indicating that it is not explained by shared finite-sample variation. The corresponding relationship is substantially weaker on MMLU, where only 4% of questions exhibit sampling instability. These results show that confidence readouts derived from the same predictive system are not empirically interchangeable and that their disagreement can provide a diagnostic of unstable sampling behavior. Confidence should therefore be treated as an explicitly defined measurement rather than as a single intrinsic scalar property of a model, particularly when it is used to inform downstream evaluation, oversight, or control.