发表机构
Prime Analytics Consulting Limited; Talbert House; Thomas More University; MAKZ(普林特分析咨询有限公司; 塔尔伯特大楼;托马斯·莫尔大学; MAKZ)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过5个专家编写的临床场景和加权评分量表,评估GPT、Claude和Gemini三个前沿模型,发现关键标准通过率低(32-42%),而低权重标准通过率高(80-90%),52%的关键标准无模型通过。
AI 中文摘要
多项选择医学基准日益饱和,最近基于评分量表的评估(如HealthBench)表明,开放式临床性能远未解决——其“困难”子集最高得分仍为32%。我们提出了一个由五个临床医生编写的临床场景组成的小型、故意困难的评估数据集,涵盖四个专业(麻醉学、内科/家庭医学、急诊医学和产科),每个场景附有一个原子化、加权、MECE的评分量表(每个任务25-62个标准;共184个标准),该量表基于临床医生起草的黄金答案编写。我们评估了三个前沿模型:GPT 5.4、Claude Opus 4.7和Gemini 3.1 Pro。平均评分通过率为0.47(Claude)、0.39(GPT)和0.37(Gemini)。核心发现是临床优先级反转:最高权重(权重5,关键)标准的通过率仅为32.4-41.7%,而低权重(权重1)标准的通过率为80-90%。108个关键(权重5)标准中有56个(52%)没有模型满足。三个LLM自动评分器在552个评分标准中,以92.8-94.7%的准确率复现了专家的是/否标签。我们将其定位为方法和初步发现贡献:这五个任务展示了一个可扩展、可辩护的流水线,可用于开发大规模基准。
英文摘要
Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from solved - its "Hard" subset top score remains 32%. We present a small, deliberately difficult evaluation dataset of five clinician-authored clinical scenarios spanning four specialties (anaesthesia, internal/family medicine, emergency medicine, and obstetrics), each accompanied by an atomic, weighted, MECE rubric (25-62 criteria per task; 184 criteria total) authored from a clinician-drafted golden answer. We evaluate three frontier models: GPT 5.4, Claude Opus 4.7, and Gemini 3.1 Pro. Mean rubric pass rates were 0.47 (Claude), 0.38 (GPT), and 0.37 (Gemini). The central finding is an inversion of clinical priority: the highest-weighted (weight-5, critical) criteria passed at only 32.4-41.7%, while low-stakes weight-1 criteria passed at 80-90%. 55 of 108 critical (weight-5) criteria (51%) were satisfied by no model. Three LLM autoraters reproduced expert met/not-met labels on 92.8-94.6% of 552 graded criteria. We position this as a methods-and-preliminary-findings contribution: the five tasks demonstrate a scalable, defensible pipeline ready to develop into a large-scale benchmark.
Comments13 pages, 4 tables