发表机构
University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对检索增强生成中的幻觉检测问题,提出证据对齐的实体验证方法,通过多维度对齐与反事实稳定性分析,在多个基准上实现持续改进。
AI 中文摘要
幻觉检测对于大型语言模型(LLMs)至关重要,因为幻觉内容在需要事实准确性的应用中造成了重大障碍。当前的检测方法主要依赖于内部信号,如不确定性和自一致性检查,利用模型的预训练知识来识别不可靠的输出。然而,预训练知识可能过时且存在覆盖限制,尤其是对于专业或近期信息。为了解决这些限制,检索增强生成(RAG)通过在推理时检索相关证据,将输出锚定在模型参数知识之外,成为一种有前景的解决方案。在本文中,我们针对一个关键且实际的学习问题——基于RAG的幻觉检测(RHD),即利用RAG通过解决信息更新挑战来增强幻觉检测。为了解决RHD,我们提出了一种新颖的方法——证据对齐的实体验证(EAEV),该方法通过利用RAG将生成的实体与检索到的证据上下文对齐,来检测实体级幻觉。具体而言,EAEV通过三个互补维度评估实体-证据对齐,并引入反事实稳定性分析,以确保在证据扰动下的稳健对齐。在多个RAG基准上的实验表明,EAEV相比现有方法取得了持续改进,并具有强大的泛化能力。
英文摘要
Hallucination detection is crucial for large language models (LLMs), as hallucinated content creates significant barriers in applications requiring factual accuracy. Current detection methods mainly depend on internal signals like uncertainty and self-consistency checks, using the model's pre-trained knowledge to identify unreliable outputs. However, pre-trained knowledge may become outdated and has coverage limitations, especially for specialized or recent information. To address these limitations, retrieval-augmented generation (RAG) has emerged as a promising solution by retrieving relevant evidence at inference time, grounding outputs beyond the model's parametric knowledge. In this paper, we target a critical and practical learning problem RAG-based hallucination detection (RHD), where RAG is employed to enhance hallucination detection by addressing information updating challenges. To address RHD, we propose a novel method Evidence-Aligned Entity Verification (EAEV), which detects entity-level hallucinations by leveraging RAG to align generated entities with retrieved evidence contexts. Specifically, EAEV evaluates entity-evidence alignment through three complementary dimensions and introduces counterfactual stability analysis to ensure robust alignments under evidence perturbations. Experiments across multiple RAG benchmarks demonstrate that EAEV achieves consistent improvements over existing methods with strong generalization capabilities.
CommentsACL 2026
DOI:10.18653/v1/2026.findings-acl.1477