发表机构
EPFL; LiGHT Laboratory(洛桑联邦理工学院; LiGHT实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究表明临床医生成对偏好并非LLM临床安全的可靠替代指标,其排名靠前的模型仍存在大量临床安全失败,提出结合偏好与安全反馈的临床调整排名方法,支持分离偏好与安全的评估实践。
AI 中文摘要
我们利用临床医生主导的平台MOOVE(Massive Open Online Validation and Evaluation,大规模开放在线验证与评估)的专家反馈,评估临床医生成对偏好是否为大型语言模型(LLM)评估中临床安全的可靠信号,该平台收集盲法成对偏好及多标准评分量表评分。临床医生在离散[-2, +2]量表上打分,负值表示临床不安全或误导性内容。基于来自28个以上国家的736余名临床医生对13个LLM输出的26804次成对判断,我们发现临床医生偏好并非安全关键性能的可靠替代指标:在成对偏好中排名靠前的模型,在“无害性”“准确性”等维度仍会出现大量具有临床意义的失败(≤-1)。这些失败在不同专科分布不均,形成了总体排名或单一数值排行榜中不可见的特定领域“禁区”。我们进一步分析了促成因素,包括提示长度、拒绝与升级行为,以及安全关键特征与表面特征的相对贡献。相当一部分偏好投票不携带任何积极安全信号,而特征分解显示,表面特征比安全关键评分量表差异能解释略多的偏好变异。最后,我们引入了结合成对偏好与评分量表衍生反馈的临床调整偏好排名,相较于单纯的Bradley-Terry强度,该排名能生成更具安全意识的排序。我们的发现支持以下评估实践:将偏好与安全分离、直接报告安全关键失败率,以及在为临床决策制定LLM排名时纳入基于临床的调整。
英文摘要
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content. Using 26{,}804 pairwise judgments across outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, we find that clinician preference is a poor proxy for safety-critical performance. Models ranking highly under pairwise preference can still exhibit substantial rates of clinically meaningful failures ($\leq -1$) on dimensions such as \emph{Harmlessness} and \emph{Accuracy}. These failures are unevenly distributed across specialties, creating domain-specific ``no-go zones'' not visible in aggregate rankings or single-number leaderboards. We further analyze contributing factors including prompt length, refusal and escalation behavior, and the relative contributions of safety-critical versus surface-level features. A substantial fraction of preference votes carry no positive safety signal, while feature decomposition shows that surface-level characteristics explain slightly more preference variation than safety-critical rubric differences. Finally, we introduce a clinically adjusted preference ranking combining pairwise preference with rubric-derived feedback, producing a more safety-aware ordering than raw Bradley--Terry strength alone. Our findings support evaluation practices that separate preference from safety, report safety-critical failure rates directly, and incorporate clinically grounded adjustments when ranking LLMs for clinical decision making.
CommentsWithdrawn by the authors because the manuscript was posted without final co-author approval