从token概率到校准置信度:数学问答的实证研究
From token probabilities to calibrated confidence: An empirical study of mathematical question answering
- RBC Borealis(加拿大皇家银行博雷利斯(RBC Borealis))
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究针对数学问答任务,探究token概率能否用于生成校准置信度,对比单遍与多遍估计器,评估多种校准方法,发现序列聚合token概率可提升置信度质量,事后校准能降低域内误差但迁移性存在非对称问题。
中文摘要 AI 辅助
大语言模型(LLM)的置信度估计旨在评估生成答案正确的概率,而校准则是让这些估计与经验准确率保持一致。已有研究表明token概率往往过于自信,我们研究这些现成的信号是否仍能为数学问答任务提供良好校准的置信度估计。我们对比了单遍估计器(复用原始生成过程中的token概率)与多遍估计器(通过验证或随机前向传播获取额外置信度信号)。尽管单个token概率可能高度饱和,但我们发现对完整序列的token概率进行聚合能捕捉到正确与错误生成之间微小但一致的差异,从而产生更具信息量的置信度估计。多遍方法可生成校准的置信度估计,我们研究了两种此类方法:通过重新提示进行自我验证(包括成本更低的原位变体),以及蒙特卡洛 dropout(从随机前向传播的变化中推导置信度)。我们还评估了两种事后校准方法:Platt缩放和等渗回归,二者均大幅降低了域内校准误差。不过,它们的数据效率随数据集难度而变化,且校准映射在不同数据集和模型间常呈非对称迁移特性。
英文摘要
Confidence estimation for large language models (LLMs) aims to estimate the probability that a generated answer is correct, while calibration aligns these estimates with empirical accuracy. Prior work has shown that token probabilities are often overconfident, we investigate whether these readily available signals can nevertheless provide well-calibrated confidence estimation for mathematical question answering. We compare single-pass estimators, which reuse token probabilities from the original generation, with multi-pass estimators, which obtain additional confidence signals through verification or stochastic forward passes. While individual token probabilities can be highly saturated, we find that aggregating token probabilities over the full sequence captures small but consistent differences between correct and incorrect generations, yielding more informative confidence estimates. Multi-pass methods can yield calibrated confidence estimates. We study two such approaches: self-verification through re-prompting, including a lower-cost in-situ variant, and Monte Carlo Dropout, which derives confidence from variation across stochastic forward passes. We further evaluate two post-hoc calibration methods, Platt scaling and isotonic regression, both of which substantially reduce in-domain calibration error. However, their data efficiency varies with dataset difficulty, and the calibration mappings often transfer asymmetrically across datasets and models.