arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向大型语言模型可靠决策的声明级置信度校准

Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models

Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang, Li Chen

arXiv 2608.22483首次发表:更新:

发表机构

Tsinghua University; Vulcan Research; AIFT; Keio Global Research Institute (KGRI); China Mobile Research Institute; Zhongguancun Laboratory(清华大学; 伏尔坎研究院; AIFT; 庆应全球研究所(KGRI); 中国移动研究院; 中关村实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大型语言模型的幻觉及置信度与事实不匹配问题,提出黑箱场景下的声明级置信度校准框架,在TriviaQA等数据集上降低了事实问题的预期校准误差。

AI 中文摘要

大型语言模型(LLMs)正越来越多地支持高风险领域的决策,但它们常出现幻觉,且表达的置信度与事实正确性不匹配。响应级置信度是一种粗糙信号:单次生成可能混合正确与错误陈述,单个数值对必须接受、拒绝或验证单个信息片段的用户而言不可操作。我们研究声明级置信度校准作为与决策相关的不确定性信号:将每个响应分解为可验证的原子声明,利用推理时的样本间一致性和自我验证信号为每个声明分配校准后的置信度。我们的框架在黑箱设置(无对数几率、无微调)下运行,直接在声明级别应用事后校准,支持对低置信度声明进行选择性干预,如证据检索或人工审核。在TriviaQA和TruthfulQA数据集上,我们评估了六个近期模型(Llama-3.1、Mistral、Qwen2.5、DeepSeek-R1、GPT-4、GPT-4o)的七个基线,结果表明,声明级分解结合事后校准可降低事实问题的预期校准误差,同时在对抗性错误前提问题上暴露失败模式,而决策者最需要可靠的不确定性估计的正是这类问题。

英文摘要

Large Language Models (LLMs) increasingly support decision-making in high-stakes domains, but they often hallucinate and express confidence that is misaligned with factual correctness. Response-level confidence is a coarse signal: a single generation can mix correct and incorrect statements, so a single number is not actionable for users that must accept, reject, or verify individual pieces of information. We study claim-level confidence calibration as a decision-relevant uncertainty signal: each response is decomposed into atomic, verifiable claims, and each claim is assigned a calibrated confidence using inference-time signals from consistency across samples and self-verification. Our framework operates in closed-box settings (no logits, no fine-tuning) and applies post-hoc calibration directly at the claim level, enabling selective intervention such as evidence retrieval or human review for low-confidence claims. Across TriviaQA and TruthfulQA we evaluate seven baselines on six recent models (Llama-3.1, Mistral, Qwen2.5, DeepSeek-R1, GPT-4, GPT-4o), and show that claim-level decomposition combined with post-hoc calibration reduces expected calibration error on factual questions while exposing failure modes on adversarial false-premise questions where decision-makers most need reliable uncertainty estimates.

CommentsIn Proceedings of The 5th Workshop on Uncertainty Reasoning and Quantification in Decision Making (held in conjunction with ACM SIGKDD 2026), Jeju, Korea

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑