arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考大语言模型中的不确定性评估

Rethinking Uncertainty Evaluation in Large Language Models

Krish Matta, Atharv Naphade, Andy Zou

arXiv 2607.19367首次发表:更新:

发表机构

Carnegie Mellon University; Meta(卡内基梅隆大学; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型中置信度评估问题,指出校准标准不充分。沿结构连贯性、忠实性和有用性三个轴形式化条件为C1指标,发现常用估计器违反条件,RLHF等未恢复连贯性,框架可测度与弥合差距。

AI 中文摘要

校准是评估大语言模型置信度的主要标准,但它并不充分:它允许明显不连贯的估计器,依赖于评估分布,并且不测试估计在多大程度上可以被解释为一致的潜在概率函数。我们实际需要的是大语言模型置信度估计满足连贯概率信念所需的条件。我们沿着三个轴(结构连贯性、忠实性和有用性)对这些条件进行形式化,并将它们作为C1指标进行操作化。尽管看起来校准良好,但广泛使用的估计器系统地违反了这些条件:模型在31%的时间里对逻辑上更容易的问题赋予较低的置信度,并且减少均方根校准误差(RMSCE)的常见干预措施并不能改变结构上的违规情况,这表明校准与概率有效性是正交的。基于人类反馈的强化学习(RLHF)和思维链并没有恢复连贯性,而是提高了有用性指标。我们的结果表明,当前大语言模型的置信度估计不能被解释为连贯概率;我们的框架提供了测量和弥合这一差距的工具。

英文摘要

Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can be interpreted as a consistent, underlying probability function. What we actually need is for LLM confidence estimates to satisfy the conditions required of coherent probabilistic beliefs. We formalize these conditions along three axes (structural coherence, faithfulness, and usefulness) and operationalize them as the C1 metrics. Widely used estimators systematically violate these conditions despite appearing well-calibrated: models assign lower confidence to logically easier questions 31\% of the time, and common interventions reducing RMSCE leave structural violations unchanged, suggesting that calibration is orthogonal to probabilistic validity. RLHF and chain-of-thought improve usefulness metrics without restoring coherence. Our results show current LLM confidence estimates cannot be interpreted as coherent probabilities; our framework provides the tools to measure and close this gap.

Comments19 pages, 11 figures, ICML 2026 EIML Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑