arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20526cs.AIcs.LGstat.ML

ConfidenceBench:评估大语言模型中的置信度校准

ConfidenceBench: Evaluating Confidence Calibration in Large Language Models

Matthew ffrench-Constant, Daniel Yang, Xinmeng Huang, Sanyam Kapoor

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型在特定场景下仅靠准确性不足的问题,提出ConfidenceBench校准基准,用Brier分数评估15个前沿LLMs口头置信度,经实验发现模型准确性与校准差异大,证明口头置信度校准是LLM可靠性重要维度,补充了基于准确性的评估。

中文摘要 AI 辅助

大语言模型(LLMs)越来越多地部署在流畅但错误答案代价高昂的场景中。在此类场景中,仅准确性不足:模型还必须知道自己何时可能出错。我们提出了ConfidenceBench,这是一个校准基准,使用Brier分数评估15个前沿LLMs中的口头置信度估计,该分数是激励真实概率报告的恰当评分规则。通过提示引出置信度,无需访问模型对数,使该框架适用于闭源和开源系统。该基准包括四类共200个私密选择题:空间推理、高精度数学、单词查找和不可知问题。在三次独立运行中,Claude Opus 4.6和Gemini 3.1 Pro Preview的Brier分数最低,均为0.103。两者均大幅优于校准随机基线0.1875,而Gemini 3.1 Flash-Lite得分为0.367,表明校准严重错误。模型家族的准确性和校准差异很大:最准确的模型并非校准最佳的,一些模型尽管准确性合理,但表现比校准随机Brier基线更差。这些结果表明,口头置信度校准是LLM可靠性一个独特且实际重要的维度,是基于标准准确性评估的补充。

英文摘要

Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone is insufficient: models must also know when they are likely to be wrong. We present ConfidenceBench, a calibration benchmark that evaluates verbalized confidence estimates in 15 frontier LLMs using the Brier score, a proper scoring rule that incentivises truthful probability reporting. Confidence is elicited via prompting, requiring no access to model logits and making the framework applicable to both closed-source and open-source systems. The benchmark comprises 200 private multiple-choice questions across four categories: spatial reasoning, high-precision mathematics, word lookup, and unknowable questions. Across three independent runs, Claude Opus 4.6 and Gemini 3.1 Pro Preview achieve the lowest Brier scores, both reported as 0.103. Both substantially outperform the calibrated-random baseline of 0.1875, while Gemini 3.1 Flash-Lite scores 0.367, indicating severe miscalibration. Accuracy and calibration diverge substantially across model families: the most accurate model is not the best-calibrated, and several models perform worse than the calibrated-random Brier baseline despite reasonable accuracy. These results show that verbalized confidence calibration is a distinct and practically important axis of LLM reliability, complementary to standard accuracy-based evaluation.

发表机构

  • ETH Zurich(苏黎世联邦理工学院)
  • University of Pennsylvania(宾夕法尼亚大学)
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

↑