arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型的分类置信度估计中的评估缺陷与稀疏性限制

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

Elena Merdjanovska, Omar Zaidan, Andreas Rücklé

arXiv 2608.04899首次发表:更新:

发表机构

Humboldt-Universität zu Berlin; Amazon(柏林洪堡大学; 亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出 LLM 分类置信度估计中 verbalization 方法存在稀疏性缺陷,AUARC 评估的插值选择会影响排名,提出 verbalization logprobs 方法可解决稀疏性并提升 AUARC 且无额外推理成本。

AI 中文摘要

当大语言模型(LLM)用于分类任务时,置信度估计至关重要,它可指示预测结果何时可信。然而,诸如 verbalization(置信度表述)之类的常用方法会产生极其稀疏的输出。例如,Qwen3-32B 在 SST-2 数据集上仅输出 8 种不同的置信度值,其中超过一半恰好为 95%,我们在 4 个数据集和 2 个 LLM 上均观察到这一一致模式。除了限制实际实用性外,我们还表明这种稀疏性会严重影响评估:accuracy-rejection 曲线下面积(AUARC)中插值方式的选择会显著改变排名,一致性采样在分段插值与线性插值之间从最佳变为最差。我们倡导将分段插值标准化以实现更公平的比较。在这种公平评估下,我们发现将 verbalization 数字按 token 概率加权的方法(我们称之为 verbalization logprobs)可解决稀疏性问题,且无需额外推理成本即可实现最佳 AUARC(较普通 verbalization 提升 2.3 个百分点)。

英文摘要

Confidence estimation is essential when LLMs are used for classification, indicating when predictions can be trusted. However, common approaches such as verbalization produce extremely sparse outputs. For instance, Qwen3-32B verbalizes only eight unique confidence values on SST-2, with over half being exactly 95%, a pattern we observe consistently across four datasets and two LLMs. Besides limiting practical utility, we show that this sparsity critically affects evaluation: the choice of interpolation in area under the accuracy-rejection curve (AUARC) dramatically alters rankings, with consistency sampling dropping from best to worst under stepwise versus linear interpolation. We advocate for standardizing stepwise interpolation for a fairer comparison. Under such a fair evaluation, we find that weighting verbalized digits by token probabilities, a method we term verbalization logprobs, addresses sparsity and achieves the best AUARC (+2.3 points over vanilla verbalization) without incurring additional inference cost.

CommentsPublished at Findings of ACL 2026

Journal refFindings of the Association for Computational Linguistics: ACL 2026, pages 33424-33435

DOI:10.18653/v1/2026.findings-acl.1671

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑