arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从 token 到语义:利用互补信号检测黑盒大语言模型的幻觉

From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs

Urja Pawar, Rajitha Ramanayake, Owen O'Neill, Nabeel Kemal, Abhishek Mandal, Houssem Chatbri, Christopher Martin, Vadim Pertsovskiy

arXiv 2609.02679首次发表:更新:

发表机构

BNY(纽约梅隆银行)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对黑盒大语言模型的幻觉检测问题,提出 TopK、CoCoA、Gated、Stacked 等方法,结合语义熵与 token 不确定性信号,在七个基准上开展评估,为不同场景提供了有效检测方案。

AI 中文摘要

当大语言模型(LLM)支持面向公众或高风险工作流程时,遗漏的虚构内容会损害用户和机构,而误报则会消耗有限的人工审核能力。当没有可信上下文或参考文档可用时,我们研究通过黑盒模型 API 可获取的两种信号:语义熵(衡量采样响应语义之间的分歧)和从 token 对数概率中衍生的不确定性。它们的失效模式具有互补性:当响应形成一个语义集群时,语义熵会变得无信息,而 token 不确定性可能会遗漏始终自信的错误。我们通过 TopK 方法扩展基于 token 的不确定性检测,该方法跨采样响应聚合 token 级信号;评估混合方法 CoCoA,其结合目标响应不确定性与语义相似度;并提出并研究两种监督方法:Gated(将单集群案例路由至聚合 token 特征分类器)和 Stacked(从语义不确定性和更广泛的 token 特征中联合学习)。我们使用四种语言模型评估七个基准,包括五个公共基准(四个文本数据集和多模态手写支票提取)以及两个构建的基准(财务摘要和长文本问答)。在我们跨模型和数据集的评估中,Stacked 在近一半的案例中表现最佳,而 TopK 和 CoCoA 在无监督训练标签的情况下仍具有竞争力,尽管它们的阈值需要仔细校准。没有一种方法是普遍最强的。因此,我们评估了假阳性率预算从 1% 到 15% 时的性能,评估它们对生成和校准选择的敏感性,并检查数据集特征的变化。

英文摘要

When LLMs support public-facing or high-stakes workflows, missed fabrications can harm users and institutions, while false alarms consume limited human-review capacity. When no trusted context or reference document is available, we study two signals accessible through black-box model APIs: semantic entropy, which measures disagreement among sampled response meanings, and uncertainty derived from token log-probabilities. Their failure modes can be complementary: semantic entropy becomes uninformative when responses form one semantic cluster, while token uncertainty can miss consistently confident errors. We extend token-based uncertainty detection by aggregating token-level signals across sampled responses through our TopK method, evaluate the hybrid CoCoA method, which combines target-response uncertainty with semantic dissimilarity, and propose and study two supervised methods: Gated, which routes single-cluster cases to an aggregated-token-feature classifier, and Stacked, which learns jointly from semantic uncertainty and broader token features. We evaluate seven benchmarks, including five public benchmarks (four text datasets and multimodal handwritten-cheque extraction) and two constructed benchmarks (Financial Summaries and Long-Text QA), using four language models. In our evaluation across models and datasets, Stacked gave the best performance in nearly half of the cases, while TopK and CoCoA remain competitive without supervised training labels, although their thresholds require careful calibration. No method is universally strongest. We therefore evaluate performance at false-positive-rate budgets from 1% to 15%, assess their sensitivity to generation and calibration choices, and examine variation across dataset characteristics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑