发表机构
University of California, Berkeley; Medical Sphere AI(加州大学伯克利分校; 医疗球体人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究应用RIFT分类法分析医学基准测试的评分标准缺陷,发现非原子性和错位问题普遍存在,捆绑标准可致分数偏差达15.9个百分点,且现有方法低估此类问题。
AI 中文摘要
医学评估正从静态的基于选项的提问转向具有开放式输出模式的现实临床场景。然而,天真地大规模评分这些输出是昂贵的,基于评分标准的评估已成为占主导地位的可扩展替代方案。我们探究当评分标准本身并非无懈可击时会发生什么,以及此类缺陷能否被检测和纠正。我们将RIFT(一种全局评分标准失败分类法)应用于两个临床基准测试(HealthBench Professional和LiveMedBench),发现失败模式具有重要意义:在HealthBench Professional上,一个LLM评判者将29.6%的标准标记为非原子性,将65.4%标记为错位/僵化。接着,我们证明这些缺陷是有实质影响的,而不仅仅是表面问题。例如,将形式为“以下至少一项/以下所有项”的捆绑标准重写为等权重的子项,并对相同响应重新评分,会使受影响对话的分数变化高达15.9个百分点,其中析取捆绑会抬高分数,而合取捆绑会压低分数。我们还发现,RIFT在临床评分标准上普遍低估捆绑问题,仅将LiveMedBench中3.3%的标准标记为非原子性,而表面形式分析在25.8%的标准中发现了结构。
英文摘要
Medical evaluation is shifting from static option-based questioning to realistic clinical scenarios with open-ended output modes. Grading these at scale naively, however, is expensive, and rubric-based evaluation has become the dominant scalable alternative. We ask what happens when the rubrics themselves are not airtight, and whether such flaws can be detected and corrected. We apply RIFT, a global rubric failure taxonomy, to two clinical benchmarks (HealthBench Professional and LiveMedBench), and find failure modes are meaningful: on HealthBench Professional an LLM judge flags 29.6% of criteria as non-atomic and 65.4% as misaligned/rigid. Then, we show that these flaws are meaningful and not simply cosmetic. As an example, rewriting bundled criteria of the form "at least one of / all of the following" as equally weighted children and regrading identical responses shifts scores by up to 15.9 percentage points on affected conversations, with disjunctive bundles inflating scores and conjunctive bundles deflating them. We also find that RIFT generally under-detects bundling on clinical rubrics, flagging 3.3% of LiveMedBench criteria as non-atomic where surface-form analysis finds structure in 25.8%.