arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KlinikeBench:超越诊断准确性的语言模型评估

KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy

Xueting Fang, Zehui Li, Yang Yang, Camilla Giovino, Shubh K. Patel, Shailly Prajapati, Vallijah Subasri, Caihua Shan

arXiv 2609.38480首次发表:更新:

发表机构

Zhejiang University; Imperial College London; Nanchang University; University of Toronto; Microsoft(浙江大学; 帝国理工学院; 南昌大学; 多伦多大学; 微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

KlinikeBench是一个由333个临床医生撰写的任务组成的基准,通过虚拟患者交互评估语言模型的完整临床诊疗能力,揭示诊断准确率(最高90.7%)与实际交互成功率(低于30%)之间的显著差距。

AI 中文摘要

大多数临床基准测试使用完整的病例描述来评估语言模型(LM)的诊断能力。然而,在临床实践中,患者以不同的方式呈现信息,临床医生必须在做出诊断之前获取相关病史并确定需要哪些检查。因此,仅凭诊断准确性无法确定智能体是否收集了必要的信息或进行了适当的临床评估。此外,现有基准缺乏专业临床医生的验证。为弥补这一空白,我们引入了KlinikeBench,一个包含333个由临床医生撰写的任务的基准,每个任务提供一个隔离的沙盒环境,其中包含虚拟患者、临床工具和特定任务的成功标准。超过35位临床医生参与了病例撰写和基准评估。在一项实证研究中,临床医生对模拟对话的平均质量评分高于参考对话(改编自真实对话)。在每个任务中,语言模型拥有固定的回合预算,用于与患者沟通、询问相关病史、请求检查、遵循行动约束并记录最终诊断。我们分别以及整体对这些步骤进行评分。在31个模型和七个模型家族中,表现最好的模型(例如GPT-6-astra和Claude Opus 5)在不到30%的任务上成功,尽管其诊断准确率达到了90.7%。一些模型受益于与患者交谈;另一些模型能从完整病历中良好诊断,但在对话中表现差得多。总体而言,KlinikeBench为评估完整的临床诊疗过程提供了一个测试平台,并揭示了诊断准确性与交互式临床评估表现之间的显著差距。

英文摘要

Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy alone therefore cannot establish whether an agent gathered essential information or conducted an appropriate clinical assessment. Furthermore, existing benchmarks lack professional clinicians' verification. To address this gap, we introduce KlinikeBench, a benchmark of 333 clinician-authored tasks, each providing an isolated sandbox environment with a virtual patient, clinical tools, and task-specific success criteria. More than 35 clinicians contributed to case authoring and benchmark evaluation. In an empirical study, clinicians gave simulated dialogues higher mean quality ratings than reference conversations, which is adapted from real conversation. In each task, an LM has a fixed budget of turns to communicate with the patient, ask about relevant history, request examinations, follow action constraints, and record a final diagnosis. We score these steps separately as well as together. Across 31 models and seven model families, the best-performing models (e.g., GPT-6-astra and Claude Opus 5) succeed on less than 30% of tasks, even though their diagnosis accuracy reaches 90.7%. Some models benefit from talking with the patient; others diagnose well from a complete chart but perform much worse in conversation. Overall, KlinikeBench provides a testbed for evaluating the full clinical encounter and reveals a substantial gap between diagnostic accuracy and performance in interactive clinical assessment. All the code and data is available on https://zehui127.github.io/klinikebench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑