临床医生对语言模型的使用方式与这些模型的评估方式存在差异
Clinician use of language models diverges from how the models are evaluated
浏览论文内容
中文总结 AI 辅助
该研究分析了6342名医护人员的12.7万余条查询,发现临床AI基准测试与实际临床使用场景的任务组合重叠度仅31%,据此提出评估应匹配实际临床使用的结论。
中文摘要 AI 辅助
大型语言模型(LLM)助手正被部署到各医疗系统的临床医生手中,对其是否可用的判断主要基于基准测试分数,其中大多数分数来自考试题目或精心整理的案例。基准测试对部署后性能的预测仅在其测试项与实际使用场景相似时才成立,但基准测试是否反映了这些系统实际承担的工作却很少被测量。本研究分析了8个月推广期间,35个专科的6342名医生、高级执业医师及护士向某机构助手发送的127833条查询。我们采用RCQ-Map(一种经临床医生验证、基于临床问题分类和LLM评估分类的框架)对每条查询进行特征刻画,记录其任务、意图、可回答性、缺失信息及潜在危害。文档处理与管理(36.2%)和知识检索(28.9%)占近三分之二的使用场景,诊断仅占3.7%;超过三分之一的查询按现有表述无法得到良好回答。将RCQ-Map应用于从主要评估套件和前沿模型报告中收集的58个公共基准测试(我们将其整合为临床AI基准图谱),结果显示:基准测试的中位数不包含文档处理请求,且任务组合与实际使用场景的重叠度仅为31%,甚至低于任务类别均匀分布的情况;旨在模拟临床实践的套件中的基准测试,与前沿模型报告中使用的基准测试相比,并未更贴近实际使用场景。因此,基准测试分数几乎无法说明临床AI在实际承担的大部分工作中的表现,评估应与实际临床使用场景相匹配。
英文摘要
Large language model (LLM) assistants are being deployed to clinicians across health systems, and judgments about their readiness rest largely on benchmark scores, most of them derived from examination questions or curated cases. A benchmark predicts performance in deployment only to the extent that its items resemble real use, yet whether benchmarks reflect the work these systems receive has rarely been measured. Here we analyze 127,833 queries sent by 6,342 physicians, advanced practice providers and nurses in 35 specialties to an institutional assistant during an eight-month roll-out. We characterize each query with RCQ-Map, a clinician-validated framework grounded in taxonomies of clinical questions and of LLM evaluation, which records its task, intent, answerability, missing information and potential harm. Documentation and administration (36.2%) and knowledge retrieval (28.9%) made up nearly two-thirds of use, and diagnosis 3.7%; more than a third of queries could not be answered well as posed. Applying RCQ-Map to 58 public benchmarks drawn from major evaluation suites and frontier model reports, which we assemble into the Clinical AI Benchmark Atlas, showed that the median benchmark contained no documentation requests and shared 31% of the task mix of real use, less than an even spread across task categories would. Benchmarks in suites designed to resemble clinical practice were individually no closer to real use than those used in frontier model reports. Benchmark scores therefore say little about how clinical AI performs on most of the work it is actually given, and evaluation should be matched to real clinical use.
发表机构
- NYU Langone Health(纽约大学朗格尼医学中心)
- Washington University School of Medicine(华盛顿大学医学院)
- New York University(纽约大学)
- Johns Hopkins University School of Medicine(约翰霍普金斯大学医学院)
- Stanford University School of Medicine(斯坦福大学医学院)
- NYU Grossman School of Medicine(纽约大学格罗斯曼医学院)
机构由 AI 辅助整理,请以论文原文为准。