概率并不足够:探索与计数分歧令牌以量化大语言模型推理不确定性
Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs
- College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机与软件学院)
- Tsinghua University(清华大学)
- Behavioral and Spatial AI Lab, Peking University & Tongji University(北京大学与同济大学行为与空间人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出分歧令牌置信度(DTC)框架,通过计数两模型解码时强烈分歧的令牌来量化LLM推理不确定性,无需训练,在数学基准上显著提升校准性能。
AI中文摘要:
随着大语言模型思维链推理能力的提升,评估和校准其推理置信度对于量化答案不确定性变得越来越重要。当前估计大语言模型置信度的方法通常基于所选关键令牌的概率,但其底层机制尚不清楚。我们的初步研究发现,用粗略替代品替换所选令牌概率也能改善校准,这促使我们进一步探索模型置信度的有效信号。我们引入了分歧令牌置信度(DTC)框架,该框架通过计数两个模型在解码过程中强烈分歧的令牌来估计置信度。DTC利用沿同一推理轨迹评估的下一令牌分布之间的Jensen-Shannon散度来识别这些分歧令牌。我们发现其计数与答案准确性几乎呈负相关,因此可作为不确定性量化的简单而有效的信号。DTC支持使用辅助模型进行白盒和黑盒评估,无需显式训练且不影响生成过程。跨多个模型家族和六个数学基准的实验表明,与基于概率和口头表达的基线相比,校准性能有所提升。在白盒评估下,仅计数估计器实现了平均期望校准误差13.0%,而标准全序列置信度方法为32.7%-42.4%。在黑盒设置中,它也改善了相对于原始口头表达分数的校准。例如,在DeepSeek-V3.2上,平均期望校准误差从32.1%-40.2%降至13.7%-16.3%。这些发现为改进大语言模型推理不确定性量化提供了新见解。代码已在此https URL发布。
英文摘要:
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen-Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby serving as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over probability-based and verbalized baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%-42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%-40.2% to 13.7%-16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at https://github.com/szu-tera/DTC.git.