arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.01103cs.CL

临床医生级别的一致性缺乏临床谨慎:LLM评估者在医学AI基准测试中的局限性

Evaluating Large Language Model Raters for German Open-Response Clinical Questions: A Physician-Annotated Benchmark Study of Agreement, Evaluator Bias, and Abstention

William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang… 展开作者

William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian Löns, Ronald Böck, Sebastian Fudickar

首次发表
浏览论文内容

中文总结 AI 辅助

针对德语开放回答临床基准MedQADE,评估LLM评估者与医生的对齐程度,发现统计一致性高但缺乏临床元认知,且存在系统性的模型谱系偏差。

中文摘要 AI 辅助

开放回答评估比多项选择基准提供更强的临床有效性,但造成了评分瓶颈,促使自动化LLM作为评判者的方法。然而,此类评估者是否复制临床校准和谨慎仍未得到检验。我们引入了MedQADE,这是第一个针对德语的标准开放回答临床基准,德语是一种缺乏本地评估基础设施的主要临床语言,包含由十名执业医师和九个大语言模型(LLM)评估者注释的3,800个项目。表现最佳的评估模型Gemini 3 Flash达到了与医生上限一致的对齐(κ = 0.694 vs. κ = 0.709),尽管宽置信区间限制了解释。尽管存在这种统计对齐,自动化评估者表现出近乎缺失的临床元认知:医生根据项目难度调整弃权,而前沿模型在每个案例中都给出确定分数。我们另外量化了系统性的谱系依赖偏差,其中模型优先评分架构兄弟,这种效应独立于语言。这些结果表明,统计对齐不能确保临床谨慎,评估者独立性需要明确验证。

英文摘要

Background: Expert-annotated benchmarks for non-English open-response clinical questions are scarce. LLM-as-a-judge systems may scale evaluation but require validation. Objective: To introduce MedQADE, a standardized German open-response clinical benchmark with physician reference annotations, and evaluate LLM-as-a-judge alignment, self- and intra-family bias, and abstention. Methods: The benchmark contains 3,800 question-answer sets with answers from five student LLMs and annotations from 10 physicians. All 10 rated the 200-question core; two primary raters assessed each of 3,600 extension questions, with the tenth resolving disagreements. Nine LLM evaluators assessed all sets. We assessed physician reliability, student-model accuracy, evaluator alignment, bias, and abstention. Results: Physicians showed moderate-to-substantial agreement on answer correctness (unweighted mean pairwise Cohen's kappa = 0.612) but limited agreement on question difficulty (Krippendorff's alpha = 0.208 using squared numeric-score distances). Student-model accuracy was 17.8%-66.0% and generally decreased with physician-rated difficulty. Gemini 3 Flash approached the leave-one-out physician reference (kappa = 0.694 vs 0.709). Four of five models rated their own responses more favorably than out-of-family evaluators; five of six intra-family comparisons were positive. Physician abstention increased with perceived difficulty. Seven of nine LLM evaluators abstained in no more than 0.51% of evaluations; the two strongest evaluators assigned definitive labels to every response. Conclusions: Strong LLM evaluators approached physician agreement, but evaluator bias and low observed abstention warrant physician validation and further assessment of selective deferral before fully automated evaluation. These results do not establish clinical safety.

发表机构

  • University of Luebeck(吕贝克大学)
  • University of Tübingen(图宾根大学)
  • University Hospital Schleswig-Holstein(石勒苏益格-荷尔斯泰因大学医院)
  • University Hospital Würzburg(维尔茨堡大学医院)
  • Charité – Universitätsmedizin Berlin(柏林夏里特医学院)
  • Genie Enterprise Deutschland GmbH

机构由 AI 辅助整理,请以论文原文为准。

↑