大型语言模型中的领域特定幻觉检测
Domain-Specific Hallucination Detection in Large Language Models
- Purdue University Fort Wayne(普渡大学韦恩堡分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出多信号幻觉检测流水线,在HaluEval上达F1=0.915,并利用DPO将生成器幻觉率降低55.9%,同时发现领域匹配预训练对跨领域适应至关重要。
AI中文摘要:
大型语言模型生成的流畅文本可能包含不忠实的陈述——这种现象被称为幻觉。我们提出了一种多信号检测流水线,结合了微调的DeBERTa-v3分类、蒙特卡洛(MC)Dropout不确定性量化以及温度缩放校准,用于响应级别的幻觉检测。在HaluEval基准上的评估中,我们的流水线在通用领域任务上达到了F1=0.915和AUROC=0.977,各任务的F1分数分别为0.97(问答)、0.96(摘要)和0.82(对话)。MC Dropout推理进一步将准确率提升至93.2%。上下文消融研究证实,模型执行的是真正的蕴含推理,而非利用表面模式,当移除知识上下文时,摘要任务的F1下降了24%。学习曲线分析显示,25%的训练数据即可捕获全数据性能的77%。在检测之外,我们将直接偏好优化(DPO)应用于Qwen2.5-0.5B生成器,将其幻觉率从85.5%降至37.7%(相对减少55.9%),该结果由我们的检测器测量。在SciFact生物医学基准上的跨领域评估显示,通用领域训练迁移效果不佳(F1=0.52),这促使我们进行领域特定的微调。在SciFact上微调的PubMedBERT达到了F1=0.63和AUROC=0.81,表明领域匹配的预训练是最强的适应策略。代码和模型可在以下网址获取:此https URL
英文摘要:
Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level hallucination detection. Evaluated on the HaluEval benchmark, our pipeline achieves F1=0.915 and AUROC=0.977 on general-domain tasks, with per-task F1 scores of 0.97 (QA), 0.96 (Summarization), and 0.82 (Dialogue). MC Dropout inference further improves accuracy to 93.2%. A context ablation study confirms the model performs genuine entailment reasoning rather than exploiting surface patterns, with summarization F1 dropping 24% when knowledge context is removed. Learning curve analysis reveals that 25% of training data captures 77% of full-data performance. Beyond detection, we apply Direct Preference Optimization (DPO) to a Qwen2.5-0.5B generator, reducing its hallucination rate from 85.5% to 37.7% (55.9% relative reduction) as measured by our detector. Cross-domain evaluation on the SciFact biomedical benchmark shows that general-domain training transfers poorly (F1=0.52), motivating domain-specific fine-tuning. PubMedBERT fine-tuned on SciFact achieves F1=0.63 and AUROC=0.81, demonstrating that domain-matched pre-training is the strongest adaptation strategy. Code and models are available at https://github.com/varunteja99/hallucination-detection-nlp