发表机构
Tarsus University; Arizona State University (ASU); Institute for Social Science Research; School of Human Evolution and Social Change; School of Computing and Augmented Intelligence (SCAI)(塔尔苏斯大学; 亚利桑那州立大学; 社会科学研究所; 人类进化与社会变革学院; 计算与增强智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 epistemic 诚实商数(EHQ)和 EHQ-3000 基准,通过行为度量评估 LLM 是否承认知识边界,发现模型间差异无法用能力解释,揭示传统评估不可见的行为差异。
AI 中文摘要
大型语言模型(LLMs)常常表现得自信、雄辩且知识渊博。一个自然的问题随之而来:它们是否知道自己不知道什么?为了回答这个问题,我们借鉴了 epistemic 诚实的概念,并开发了一种新的度量方法,以系统评估 LLM 是否恰当地承认其知识的边界。在这项工作中,我们引入了 epistemic 诚实商数(EHQ),它报告了跨两个操作轴(epistemic 克制和实质性回答校准)的三个可观察子分数,并构建了 EHQ-3000,一个包含 3,000 个问题的基准,涵盖虚构实体、截止后事件、超小众真实和上下文条件问题。从 21 个模型 API 路线的冻结注册表中,15 个在端点和资格检查后完成了协议;14 个进入了确认性分析,因为严重的提供方端截断使一个路线的分数无法确定。该研究揭示了模型之间的显著差异,包括一种无法通过它们提取明确可用信息的能力来解释的差异。在分析的面板中,综合 EHQ 范围从 0.31 到 0.81,尽管在基于文档的能力探针上表现接近上限。在当前类别组成下,两个克制标准强烈重叠,而实质性回答校准在不同模型间变化,且不与克制可靠地共变;然而,小面板留下了相当大的不确定性。因此,EHQ 揭示了传统基于正确性的评估所不可见的行为差异,同时也显示了为什么数据集组成、提供方行为和置信度引导必须仍然是解释的一部分。
英文摘要
Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question arises: do they know what they don't know? To answer this question, we borrow the concept of epistemic honesty and develop a novel metric to systematically evaluate whether an LLM appropriately acknowledges the boundaries of its knowledge. In this work, we introduce the Epistemic Honesty Quotient (EHQ), which reports three observable sub-scores across two operational axes (epistemic restraint and substantive-answer calibration), and construct EHQ-3000, a 3,000-question benchmark spanning Fabricated Entity, Post-Cutoff Event, Hyper-Niche True, and Context-Conditioned Questions. From a frozen registry of 21 model API routes, 15 completed the protocol after endpoint and eligibility checks; 14 entered the confirmatory analysis because severe provider-side truncation made one route's score indeterminate. The study reveals substantial variation across models, including a difference that can not be explained by their capability to extract explicitly available information. Composite EHQ ranges from 0.31 to 0.81 across the analysed panel, despite near-ceiling performance on the document-grounded capability probe. The two restraint criteria overlap strongly under the present category composition, whereas substantive-answer calibration varies across models and does not reliably co-vary with restraint; however, the small panel leaves substantial uncertainty. Thus, EHQ reveals behavioral differences that are not visible to conventional correctness-based assessment, while also showing why dataset composition, provider behavior, and confidence elicitation must remain part of the interpretation.