发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过干预训练与上下文数据中的不确定性来源,发现大语言模型的言语化概率与内部概率均受分布性和断言性不确定性影响,且两者对齐程度超预期,表明言语化概率可探测模型内部分布。
AI 中文摘要
大语言模型在其采样分布中携带一种内部的不确定性概念,即模型在生成一个答案而非另一个答案时所赋予的概率。模型也可以被要求用语言或数字陈述其置信度,即言语化的不确定性。先前的研究表明,内部概率追踪训练数据中的相对频率,而言语化概率追踪训练数据中的显式概率断言。然而,我们尚不清楚这两种读出是否对齐,除非训练数据中的频率与概率断言恰好一致。这限制了我们对于何时可以将言语化不确定性用作训练数据频率或模型内部分布的代理的理解。我们通过系统地探索训练数据与上下文数据中潜在的不确定性来源进行干预,从而填补了这一空白,考察了这些干预如何影响大语言模型的概率读出。我们发现,内部概率读出和言语化概率读出均受到训练数据中分布性不确定性和断言性不确定性的影响。此外,我们发现言语化概率与内部概率的对齐程度超出了独立追踪相同不确定性来源所预期的水平,这表明言语化概率可用于探测模型的内部分布。
英文摘要
Large language models carry an internal notion of uncertainty in their sampling distribution, i.e., the probabilities they place on generating one answer rather than another. They can also be asked to state a confidence, in words or as a number: a verbalized uncertainty. Prior work suggests that internal probabilities track relative frequencies in the training data, and that verbalized probabilities track explicit probabilistic assertions in the training data. However, we do not know whether these two readouts are aligned, except when frequencies and probabilistic assertions in the training data happen to align. This limits our understanding of when we can use verbalized uncertainties as a proxy for either training data frequencies, or a model's internal distribution. We resolve this gap by systematically exploring how LLMs probability readouts are impacted by training and in-context data, via intervening on the underlying uncertainty sources in the data. We find that both internal and verbalized probability readouts are impacted by both distributional and asserted uncertainty in the training data. Further, we find that verbalized and internal probabilities are aligned beyond what would be expected by independently tracking the same uncertainty sources, suggesting that verbalized probabilities can be used to probe a model's internal distribution.