arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用转移解码检测大语言模型中的幻觉

Hallucination Detection in Large Language Models Using Diversion Decoding

Basel Abdeen, S M Tahmid Siddiqui, Meah Tahmeed Ahmed, Anoop Singhal, Latifur Khan, Punya Parag Modi, Ehab Al-Shaer

arXiv 2607.10476首次发表:更新:

发表机构

The University of Texas at Dallas; National Institute of Standards and Technology; Carnegie Mellon University(德克萨斯大学达拉斯分校; 美国国家标准与技术研究院; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型中的幻觉检测问题,提出转移解码新方法,通过在解码阶段挑战模型响应提取特征,训练机器学习模型度量不确定性,该方法计算复杂度低,优于现有方法。

AI 中文摘要

大语言模型已成为通过无缝、类人交互检索知识的强大工具。尽管其具有先进文本生成能力,但存在幻觉倾向,产生事实错误陈述和编造知识,损害其可靠性和可信度。多项研究探索评估大语言模型不确定性和检测幻觉的方法,但现有方法通常是概率性的且计算成本高,限制了实际应用。本文引入转移解码,一种在解码阶段通过主动挑战模型生成的响应来开发大语言模型不确定性启发式方法的新方法。通过转移解码提取捕捉大语言模型产生替代答案抗性的特征,并利用这些特征训练机器学习模型以开发大语言模型不确定性的启发式度量。实验结果表明转移解码以显著更低的计算复杂度优于现有方法,是评估幻觉检测的高效且稳健的解决方案。

英文摘要

Large language models (LLMs) have emerged as a powerful tool for retrieving knowledge through seamless, human-like interactions. Despite their advanced text generation capabilities, LLMs exhibit hallucination tendencies, where they generate factually incorrect statements and fabricate knowledge, undermining their reliability and trustworthiness. Multiple studies have explored methods to evaluate LLM uncertainty and detect hallucinations. However, existing approaches are often probabilistic and computationally expensive, limiting their practical applicability. In this paper, we introduce diversion decoding, a novel method for developing an LLM uncertainty heuristic by actively challenging model-generated responses during the decoding phase. Through diversion decoding, we extract features that capture the LLM's resistance to produce alternative answers and utilize these features to train a machine-learning model to develop a heuristic measure of the LLM's uncertainty. Our experimental results demonstrate that diversion decoding outperforms existing methods with significantly lower computational complexity, making it an efficient and robust solution for evaluating hallucination detection.

Journal refData and Applications Security and Privacy XXXIX. DBSec 2025

DOI:10.1007/978-3-031-96590-6_7

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑