arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08026cs.CL

幻觉检测基准中的标注问题:一项实证评估

The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation

Jorma Valjakka, Juhani Kivimäki, Juha Mylläri, Jukka K. Nurminen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过900个人工标注问答对实证评估了幻觉检测基准中参考忠实性与事实正确性标准不匹配的问题,发现自动化标注策略间及与人工标注存在显著分歧,提示设计对结果影响大,标注来源选择应成为基准设计的关键部分。

中文摘要 AI 辅助

近年来,针对大型语言模型(LLM)产生幻觉的检测方法已开发出多种。这些方法通常使用包含问题及相应简短参考答案的开放域问答(QA)数据集进行基准测试。首先,使用LLM生成对QA数据集中问题的答案。然后,采用某种自动化标注策略,通过将答案与数据集中的参考答案进行比较,来标注这些答案是否为幻觉。这种评估设置在两个标准之间造成了方法论上的模糊性:参考忠实性(答案是否完全由参考支持)和事实正确性(答案是否没有矛盾且不包含事实错误的具体主张)。在实践中,即使预期目标是后者,自动化标注器也可能应用前一个标准。我们使用900个由人工标注的问答对来研究这种潜在的标准不匹配,这些问答对跨越三个常用的QA数据集和三个生成模型,标注目标为答案级的事实正确性。我们在受控提示变体下评估了词汇相似性度量、一个参考蕴含的NLI基线以及七个LLM评判器作为自动化标注器。我们的实验揭示了自动化标注策略之间以及这些标注与人工标注之间的显著分歧。许多策略还表现出强烈的方向性错误偏差,并且对于大多数评判器-生成器组合,将面向忠实性的提示替换为面向事实正确性的提示,可提高与人工标注的一致性并减少假阳性主导,表明自动化幻觉标注在很大程度上取决于目标标准的具体指定方式。因此,标注来源的选择应被视为基准设计的基本组成部分,并应明确、验证且与基准目标相匹配。

英文摘要

In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness (whether the answer is fully supported by the reference) and factual correctness (whether the answer is free from contradictions and factually false specific claims). In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question-answer pairs spanning three commonly used QA datasets and three generator models, with labels targeting answer-level factual correctness. We evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants as automated labelers. Our experiments reveal substantial disagreement both among automated labeling strategies and between these labels and human annotations. Many strategies also exhibit strong directional error biases, and for most judge-generator pairs, replacing a faithfulness-oriented prompt with a factual-correctness prompt improves agreement with human annotations and reduces false-positive dominance, indicating that automated hallucination labels depend strongly on how the target criterion is specified. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.

发表机构

  • University of Helsinki(赫尔辛基大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑