arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12718cs.AI

当评分标准失效:幻觉揭示医学AI评估中的盲区

When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation

发表机构牛津大学
查看机构详情
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Griffin Farrow, Lily Sijia Li, Jack Johnson, Tingyan Wang, Philip Torr, William Bolton, Fabio J. Fehr

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示基于评分标准的医学LLM评估存在系统性盲区,无法捕捉临床相关幻觉,提出分类法与错误注入流水线验证,并建议结合检索式事实核查以补充评估。

中文摘要 AI 辅助

幻觉可能削弱临床医生对大型语言模型(LLM)的信任,因此评估方法必须能够捕捉临床相关的错误。基于评分标准(rubric)的评估已成为医学领域评估LLM的主要方法,但尚不清楚评分分数是否能反映此类错误。我们首先在受控环境中使用MedHallu数据集进行研究,发现更具体的评分标准能更好地区分正确回答与幻觉回答。为系统验证这一点,我们开发了一套医学幻觉类型分类法,以及一个经临床医生验证的错误注入流水线,该流水线能生成匹配的正确回答和注入错误的回答。在HealthBench、HealthBench Professional和LiveMedBench基准上,我们发现的临床相关幻觉未被评分标准捕捉,分数往往保持不变。我们发现,评分标准在明确核查事实时最为有效,而对于其未预见的额外或意外错误则效果较差。一种基于检索的事实性初步检查能找回部分评分标准盲区中的错误,表明这是一种互补方法。这些发现揭示了当前LLM医学评估中系统性的盲区,并表明仅凭评分标准分数不足以确立临床可靠性,可能削弱临床医生对临床部署的信任和信心。

英文摘要

Hallucinations can undermine clinician trust in LLMs, making it important that evaluation methods capture clinically relevant errors. Rubric-based evaluation has become the leading approach for assessing LLMs in medicine, but it is unclear whether rubric scores reflect such errors. We first study this in a controlled setting using MedHallu, finding that more specific rubrics better distinguish correct from hallucinated responses. To test this systematically, we develop a taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that creates matched correct and error-injected responses. Across HealthBench, HealthBench Professional, and LiveMedBench, our clinically relevant hallucinations are missed by rubrics, often leaving scores unchanged. We find that rubrics are most effective when explicitly checking facts, and are less effective for additional or unexpected errors they do not anticipate. A preliminary retrieval-based factuality check recovers some of the rubric-blind errors, suggesting a complementary approach. These findings reveal systematic blind spots in current medical evaluation of LLMs and suggest that rubric scores alone are insufficient to establish clinical reliability, potentially undermining clinician trust and confidence in clinical deployment.

↑