发表机构
University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过元分析和问答实验,揭示自动化评判者在评估不确定性量化器时,其错误对评估可靠性的影响难以预测,且易误导信息量大的UQ,并关联了置信度与错误类型的模式。
AI 中文摘要
LLM在广泛的NLG应用中的广泛采用,凸显了为用户提供避免错误和幻觉的手段的重要性。不确定性量化有望填补这一空白;低不确定性(高置信度)作为正确性的代理,允许用户进行选择性操作(例如,拒绝低置信度、可能不正确的响应)。置信度与正确性之间的相关性因此成为评估不确定性量化器(UQs)的有用标准。但在NLG中,由于多种响应可能都适用于同一提示,获得可靠的正确性判断并不简单,尤其是在没有人工干预的情况下。自动化判断中的错误几乎不可避免,并且已知会降低评估协议的可信度(Santilli等人,2025;Ielanskyi等人,2025)。在对已发表工作的元分析中,我们表明自动化判断是当前的主流。此外,自动化判断很少与人工判断进行验证,而对其所自动化的UQ评估的验证则更为罕见。通过在问答实验中使用4个LLM、人工和自动化判断以及7种流行的UQ,我们发现:i)评判者的表现只能粗略预测其错误对UQ评估可靠性的实际影响;ii)判断错误往往最容易误导信息量大的UQ。我们将这些观察与置信度和判断错误类别之间的相关性模式联系起来。
英文摘要
The wide adoption of LLMs across broad NLG applications heightens the importance of providing users with the means to avert errors and hallucinations. Uncertainty quantification is poised to fill that gap; with low uncertainty (high confidence), as a proxy for correctness, allowing users to be selective (e.g., reject low-confidence, likely incorrect responses). Correlation between confidence and correctness then serves as a useful criterion for evaluation of uncertainty quantifiers (UQs). But in NLG, where diverse responses can be adequate to a prompt, obtaining reliable correctness judgements is not simple, especially without human intervention. Errors in automated judgement are hardly avoidable and known to diminish the reliability of evaluation protocols (Santilli et al., 2025; Ielanskyi et al., 2025). In a meta-analysis of published work, we show that automated judgement is the present norm. Besides, automated judgements are rarely validated against human ones, and the validation of the UQ evaluation they automate is even rarer. With experiments in question answering, using 4 LLMs, human and automated judgements and 7 popular UQs, we find that i) a judge's performance can only coarsely predict the observed impact of its errors on the reliability of UQ evaluation, and that ii) judgement errors tend to misrepresent informative UQs most. We link these observations to patterns of correlation between confidence and categories of judgement error.
CommentsAccepted at INLG 2026