arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

核令牌矛盾:一种快速且原则性的大语言模型声明不确定性量化方法

Kernel Token Contradiction: a Fast and Principled Approach for LLM Claim Uncertainty Quantification

Jérémie Dentan, Alexi Canesse, Mahammed El Sharkawy, Sonia Vanier

arXiv 2608.22506首次发表:更新:

发表机构

LIX (École Polytechnique, IP Paris, CNRS)(LIX(巴黎综合理工学院、巴黎IP大学、法国国家科学研究中心))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出KTC方法,实现仅CPU运行下8.2倍于GPU交叉编码器方法、65倍于同性能CPU方法的加速,在多语言多模型基准上匹配或超越现有声明级UQ方法,可用于LLM输出实时监控。

AI 中文摘要

声明级不确定性量化(UQ)旨在通过评估大语言模型(LLM)输出中每个声明的真实性,缓解LLM可靠性不足的问题。我们提出核令牌矛盾(KTC),一种在实际白盒条件下计算声明级UQ的轻量方法。KTC将LLM生成涉及的候选令牌表示为半正定核,该核整合了LLM的条件分布和令牌矛盾得分。随后,我们使用冯·诺依曼熵量化该核的不确定性。为估计令牌矛盾,我们开发了一种基于维基百科语料库频率统计的新方法。尽管仅使用CPU,我们的方法与基于交叉编码器的最先进GPU加速方法相比,实现了超过8.2倍的加速,与性能相当的仅CPU方法相比,实现了超过65倍的加速。我们的评估覆盖了两个基准、四种欧洲语言和16种不同模型。KTC不仅达到了现有方法的平均性能,还在高精度 regime 中表现优于它们。这种计算效率与准确性的结合,使得在生产环境中对LLM输出进行实时监控成为可能。

英文摘要

Claim-level Uncertainty Quantification (UQ) aims to mitigate the lack of reliability of Large Language Models (LLMs) by evaluating the factuality of each claim in their outputs. We introduce Kernel Token Contradiction (KTC), a lightweight approach to compute claim-level UQ under realistic white-box conditions. KTC represents the candidate tokens involved in LLM generation as a positive semi-definite kernel that integrates both the LLM's conditional distribution and a token contradiction score. We then use the Von Neumann entropy to quantify the uncertainty of this kernel. To estimate token contradiction, we develop a new approach based on frequency statistics from the Wikipedia corpus. Although CPU-only, our approach achieves over an 8.2x speedup compared to state-of-the-art GPU-accelerated methods based on cross-encoders, and over a 65x speedup compared to CPU-only methods with comparable performance. Our evaluation spans two benchmarks across four European languages and 16 different models. KTC not only matches the average performance of existing methods but also outperforms them in high-precision regimes. This combination of computational efficiency and accuracy makes real-time monitoring of LLM outputs practical in production.

CommentsPreprint. Under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑