arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当监督式不确定性量化集成能提升大语言模型幻觉检测性能?一项鲁棒性研究

When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study

Mohit Singh Chauhan, Vipin Gyanchandani, Dylan Bouchard

arXiv 2608.24492首次发表:更新:

发表机构

CVS Health(CVS健康公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究监督式UQ集成对LLM幻觉检测的鲁棒性,在多模型、多数据集、多生成范式下分析其性能,发现其多数场景优于单个评分器,采样黑盒集成效果接近全集成。

AI 中文摘要

不确定性量化(UQ)方法广泛应用于闭卷场景下大语言模型(LLM)的幻觉检测,该场景在推理时无法获取真实证据。已有研究提出通过学习集成结合UQ信号,但对这类集成的鲁棒性实证研究有限。本文研究一种监督式集成框架:在小型领域特定的标注LLM响应数据集上,对异构的基于UQ的评分器输出训练分类器,随后将其应用于无检索、无工具、无参考文档的样本外幻觉分类。针对4种LLM、9个数据集和3种生成范式(短格式问答、长格式生成、代码生成),本文从样本效率、域内数据集迁移、生成范式依赖性三个维度开展系统鲁棒性分析。研究发现,在32项设置中,监督式集成在30项上优于最优单个评分器,仅需100个标注实例即可实现性能提升;在分布偏移下的域内迁移场景中,集成保留了大部分优势,在28项迁移设置中,23项优于最优非集成评分器;基于采样的黑盒集成性能与全集成相近,而单生成白盒集成的收益有限。

英文摘要

Uncertainty quantification (UQ) methods are widely used for hallucination detection in large language models (LLMs) in closed-book settings where ground-truth evidence is unavailable at inference time. Prior work has proposed combining UQ signals via learned ensembles, but empirical investigations into the robustness of these ensembles are limited. We study a supervised ensembling framework that trains a classifier over heterogeneous UQ-based scorer outputs on a small, domain-specific dataset of labeled LLM responses, then applies it to out-of-sample hallucination classification without retrieval, tools, or reference documents. Across four LLMs, nine datasets, and three generation regimes (short-form QA, long-form generation, and code generation), we provide a systematic robustness analysis along three axes: sample efficiency, in-domain dataset transfer, and generation regime dependence. We find that supervised ensembles outperform the best individual scorer in 30 of 32 settings, with gains realized from as few as 100 labeled instances. Ensembles retain most of their advantage in cases of in-domain transfer under distribution shift, outperforming the best non-ensemble scorer in 23 of 28 transfer settings. Sampling-based black-box ensembles are nearly as effective as full ensembles, while single-generation white-box ensembles offer limited benefit.

CommentsAccepted to AACL 2026 (Findings)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑