arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

“非常可能”意味着“不确定”吗?LLM在语言不确定性量化中与人类的差异

"very likely" Means "uncertain"? How LLMs Diverge from Humans in Linguistic Uncertainty Quantification

Jinhao Duan, Zicheng Liu, Zijie Liu, Kaidi Xu, Tianlong Chen

arXiv 2610.00083首次发表:更新:

发表机构

UNC-Chapel Hill; The University of Hong Kong; City University of Hong Kong(北卡罗来纳大学教堂山分校; 香港大学; 香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究LLM与人类在语言不确定性量化上的差异,提出优化算法METHODNAME,直接从LLM输出学习标记-不确定性映射,实现标记级置信度语义比较,揭示系统性差异。

AI 中文摘要

人类通过语言标记(如“可能”、“很可能”)来表达不确定性,然而大多数大语言模型(LLM)的不确定性量化(UQ)依赖于基于似然性或一致性的信号。从认知角度看,准确的语言不确定性反映了元认知监控,代表了知识边界(“知道自己不知道”),以支持调节和信息寻求。在本文中,我们研究了LLM在语言不确定性量化中如何与人类产生分歧,以及语言标记能否可靠地量化LLM的不确定性。我们从心理学和决策科学文献中整理了一个人类不确定性标记语料库,并以此为基准对LLM进行评测。我们观察到,LLM用数值水平编码语言不确定性,这些数值水平与人类的数值水平存在显著差异。然后,我们引入了METHODNAME,一种新颖的基于优化的算法,该算法直接从LLM输出中学习不确定性标记上的最优不确定性分布。通过拟合一个标记-不确定性映射以最好地解释经验正确性,METHODNAME发现每个语言标记应传达多少概率质量,而不是通过重复采样来估计不确定性。METHODNAME实现了人类与LLM之间在置信度语义上的直接、标记级比较,解开不匹配,并揭示语言表达中系统性的置信度差异。

英文摘要

Humans express uncertainty verbally via markers (e.g., "possible," "likely"), yet most LLM uncertainty quantification (UQ) relies on costing likelihood- or consistency-based signals. From a cognitive perspective, accurate verbal uncertainty reflects metacognitive monitoring, representing knowledge boundaries ("knowing that you don't know") to support regulation and information seeking. In this paper, we investigate how LLMs diverge from humans in verbal uncertainty quantification and whether verbal markers can reliably quantify LLM uncertainty. We curate a corpus of human uncertainty markers from psychology and decision-science literature and benchmark LLMs against it. We observe that LLMs encode verbal uncertainty with numerical levels that differ substantially from those of humans. We then introduce METHODNAME, a novel optimization-based algorithm that learns an optimal uncertainty profile over uncertainty markers directly from LLM outputs. By fitting a marker-uncertainty mapping to best explain empirical correctness, METHODNAME discovers how much probability mass each verbal marker should convey, rather than estimating uncertainty via repeated sampling. METHODNAME enables a direct, marker-level comparison of confidence semantics between humans and LLMs, disentangling mismatch and revealing systematic confidence disparities in verbal expressions.

CommentsICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑