arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于不确定性下语义承诺的Credal大语言模型

Credal Large Language Models for Semantic Commitment under Uncertainty

Shireen Kudukkil Manchingal, Sofiia Nikolenko, Fabio Cuzzolin

arXiv 2608.23244首次发表:更新:

发表机构

Oxford Dynamics; Ludwig-Maximilians-Universität München; Institute for Artificial Intelligence, Data Analysis and Systems (AIDAS); School of Engineering Computing & Mathematics; Oxford Brookes University(牛津动力学; 慕尼黑大学; 人工智能、数据分析与系统研究所; 工程计算与数学学院; 牛津布鲁克斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Credal大语言模型(CLLMs),通过LoRA适配器集成体构建Credal集合,推导CTC和SCC两类承诺分数,在多模型多数据集上验证其在问答、幻觉检测等任务中表现优异。

AI 中文摘要

大语言模型(LLMs)常产生流畅但错误的答案,且置信度毫无根据。其核心局限在于,标准LLM通过单一预测分布表示不确定性,将认知无知与真实歧义混为一谈。我们提出Credal大语言模型(CLLMs):由LoRA适配器组成的集成体诱导出一个Credal集合,其上下概率可揭示合理预测分布的范围,而非坍缩为单一的softmax输出。基于该表示,我们推导了两个互补的承诺分数:Credal Token Commitment(CTC)是一种令牌空间分数,结合了下界支持、Credal宽度和交集熵,无需额外生成即可计算;Semantic Commitment Consistency(SCC)将承诺扩展到语义空间,通过采样补全实现,其中SCC-Gap用于衡量令牌级与语义级支持之间的不匹配。我们在Gemma-2-9B、Llama-3.1-8B和Qwen2.5-7B模型上,针对OpenBookQA、CoQA、TriviaQA和ARC-Challenge数据集,对幻觉检测、校准、选择性预测和推理进行了评估。CLLM在问答准确率上表现最佳,且具有竞争力的预期校准误差;CTC在无需额外生成的情况下,在多数设置中跟踪最佳幻觉AUROC,差距不超过1.5个百分点。在80%覆盖率的选择性预测任务中,带SCC的CLLM在OpenBookQA上达到99.0%的准确率;在ARC-Challenge上,带Csem置信度的CLLM在三个主干模型上实现了不超过0.6%的预期校准误差(ECE)。

英文摘要

Large language models (LLMs) often produce fluent but incorrect answers with unwarranted confidence. A central limitation is that standard LLMs represent uncertainty through a single predictive distribution, conflating epistemic ignorance with genuine ambiguity. We introduce Credal Large Language Models (CLLMs): an ensemble of LoRA adapters induces a credal set whose lower and upper probabilities expose the spread of plausible predictive distributions rather than collapsing to a single softmax output. From this representation, we derive a single commitment rule: the model commits to an answer only when its lower probability exceeds the upper probability of every alternative, and otherwise returns the set of answers that no plausible predictor rules out. We apply this commitment rule at two depths: Credal Token Commitment (CTC) applies it to answer tokens from one ensemble forward pass, which decides constrained answers without any generation; for open-ended answers, credal decoding extends a partial answer only when no completed answer dominates it, so that the completions produced are those the plausible predictors license, and Credal Semantic Commitment (CSC) applies the rule to their meaning clusters. We evaluate CLLMs with Gemma-2-9B, Llama-3.1-8B and Qwen2.5-7B on OpenBookQA, CoQA, TriviaQA and ARC-Challenge. On multiple choice, CTC commits on 73-91% of questions at 89-98% accuracy, returns sets of 1.1-1.5 options containing the gold one on 89-98%, and its intervals contain the observed accuracy in 24 of 30 confidence bins without calibration; corrupted context lowers commitment from 87-92% to 65-71%, and on Gemma the credal bound detects corruption better than every baseline. On open-ended QA, CLLM outperforms semantic entropy and Laplace-LoRA at a fixed coverage by up to 19% and 9.5% absolute accuracy on CoQA and TriviaQA with context, for every backbone.

Comments45 pages, 10 figures, 19 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑