arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越多数投票:面向医学幻觉检测的多视角裁决

Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection

Joe Cecil, Marjorie Freedman

arXiv 2609.03953首次发表:更新:

发表机构

Information Sciences Institute; University of Southern California(信息科学研究所; 南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对医学聊天机器人响应开展多视角标注研究,结合首轮标注、LLM-as-a-Judge候选发现及医学专家与证据裁决,发现单轮幻觉基准易低估事实错误,多轮裁决可提升覆盖范围。

AI 中文摘要

了解聊天机器人生成文本中事实错误的频率并评估检测这些错误的系统,对于确定聊天机器人的安全性至关重要。然而,事实错误检测通常被视为一次性的单标注者标注问题。在长篇聊天机器人响应中,事实错误可能很微妙,且嵌入在大部分正确的文本中。我们针对医学相关的聊天机器人响应开展多视角标注研究,结合首轮标注、大语言模型作为评判者(LLM-as-a-Judge,LaJ)候选发现,以及两种形式的裁决:医学专家裁决和基于证据的事实核查。首轮标注者经常遗漏后续被裁决者验证的事实错误。LaJ可提升候选发现效果,但单独使用并不足够:它会遗漏标注者能捕捉到的事实错误。我们还发现裁决者之间存在分歧,这表明整合多个候选来源的裁决可提升基准的完整性,但无法消除运用判断和专业知识的必要性。将该技术应用于现有基准时,其揭示出类似的标注遗漏模式。综合来看,这些结果表明,在所考察的场景中,一次性幻觉基准可能以低估事实错误数量为代价实现规模化。多轮裁决可提升覆盖范围,但从基准中得出的推断仍对用于确定错误存在与否的判断、专业知识和证据敏感。

英文摘要

Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-error detection is often treated as a single-pass, single-annotator labeling problem. In long-form chatbot responses, factual errors can be subtle and embedded within mostly correct text. We develop a multi-perspective annotation study of medically relevant chatbot responses, combining first-pass annotation, LLM-as-a-Judge (LaJ) candidate discovery, and two forms of adjudication: medical-expert and evidence-based fact-checking. First-pass annotators frequently miss factual errors later validated by adjudicators. LaJ improves candidate discovery, but is insufficient on its own: It misses factual errors that annotators catch. We also find disagreement among adjudicators, suggesting that adjudication over multiple candidate sources can improve benchmark completeness, but does not eliminate the need to apply judgment and expertise. Applied to an existing benchmark, this technique reveals a similar pattern of missing annotations. Together, these results suggest that in the settings examined here, single-pass hallucination benchmarks may achieve scale at the cost of undercounting factual errors. Multi-pass adjudication can improve coverage, but inferences drawn from the benchmarks are still sensitive to the judgment, expertise, and evidence used to determine error presence.

Comments34 pages, 6 figures, to be published in Findings of the ACL: EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑