arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语义不确定性量化需要事实等价性

Semantic Uncertainty Quantification Needs Factual Equivalence

Joseph Hoche, Quentin Guimard, Gianni Franchi

arXiv 2609.34967首次发表:更新:

发表机构

AMIAD(AMIAD)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型语义不确定性量化中现成算子无法准确衡量事实等价性的瓶颈,提出用对比学习训练的单一编码器作为算子,在126个评估设置中95%提升性能,最佳AUROC达0.76。

AI 中文摘要

大型语言模型的语义不确定性量化依赖于一个常见模板:抽取多个答案,衡量它们之间的一致性,并将不一致视为不确定性。我们首先将此模板形式化为两个独立角色:一个比较两个答案的算子,以及一个将所有成对比较聚合为标量的聚合器。现有方法在聚合方式上几乎完全不同,而算子则直接采用现成的,通常是NLI模型或通用句子编码器。我们表明,这种对现成算子的依赖是语义不确定性量化的主要瓶颈:它们不能准确衡量同一问题多个答案的事实等价性。我们用一个刻意简单的方案解决这个问题:一个通过对比学习训练以隔离目标事实的单一编码器,利用由LLM生成的合成数据和与所有评估设置不相交的数据集。将所得算子整合到现有方法中,在跨越语言和视觉-语言模型的18个模型-数据集组合的126个评估设置中的120个(95%)上提升了性能。最佳变体达到0.76的平均AUROC,而最强基线为0.68,同时将基于蕴含的算子的二次交叉编码器比较替换为每个答案一次编码器传递。改进的一致性支持了算子而非聚合器是限制因素的观点。同一算子还改进了单次生成token级估计器:它为每个token分配的范数衡量该token对答案的贡献程度,相应地重新加权token对数似然可锐化估计。

英文摘要

Semantic uncertainty quantification for large language models rests on a common template: sample several answers, measure how much they agree, and treat disagreement as uncertainty. We first formalize this template as two separate roles: an operator that compares two answers, and an aggregator that combines all pairwise comparisons into a scalar. Existing methods differ almost entirely in how they aggregate, while taking the operator off the shelf, typically an NLI model or a generic sentence encoder. We show that this reliance on off-the-shelf operators is the primary bottleneck of semantic UQ: they do not accurately measure factual equivalence of multiple answers to the same question. We resolve this with a deliberately simple recipe: a single encoder trained contrastively to isolate the targeted fact, utilizing synthetic data generated by an LLM and dataset both disjoint from all evaluation settings. Integrating the resulting operator into existing methods improves performance on 120 of 126 evaluation settings (95%) spanning 18 model dataset combinations across language and vision-language models. The best variant reaches 0.76 mean AUROC against 0.68 for the strongest baseline, while replacing the quadratic cross-encoder comparisons of entailment-based operators with one encoder pass per answer. The uniformity of the improvement supports the view that the operator, not the aggregator, is the limiting factor. The same operator also improves single generation token-level estimators: the norm it assigns to each token measures how much that token bears on the answer, and reweighting token log-likelihoods accordingly sharpens the estimate.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑