arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将临床判断扩展到医学AI评估

Scaling Clinical Judgment to Evaluate Medical AI

Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Jason A. Freed, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman, Arjun K. Manrai

arXiv 2609.12822首次发表:更新:

发表机构

Harvard Medical School; Beth Israel Deaconess Medical Center; Stanford University; University of Alberta; Massachusetts General Hospital; Erasmus Medical Center; University of Maryland School of Medicine; University of Maryland Institute for Health Computing; VA Maryland Healthcare System; Brigham and Women’s Hospital(哈佛医学院; 贝斯以色列女执事医疗中心; 斯坦福大学; 阿尔伯塔大学; 麻省总医院; 伊拉斯姆斯医学中心; 马里兰大学医学院; 马里兰大学健康计算研究所; 弗吉尼亚州马里兰医疗系统; 布莱根妇女医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出PrecepTron,一个经微调的LLM,用于医生级评估医学AI回答,并发布GRAND-ROUNDS基准,以可扩展方式重现多项临床研究结论,并探索LLM在医学中的推理机制。

AI 中文摘要

盲法医生评估一直被许多人视为评估大型语言模型(LLMs)临床推理能力的金标准。然而,这种方法难以扩展;因此,以往的研究通常依赖于小规模的医生评审团,且这些评审团往往来自单一机构或专科,这既限制了所研究的科学问题,也使得研究结果是否能在不同的评审者群体中重现变得不明确。为了更严谨且可扩展地研究AI模型中的临床推理,我们在此引入PrecepTron,一个经过微调用于对开放式回答进行医生级评估的LLM。PrecepTron通过低秩适配(LoRA)技术,在少量医生示例上对320亿参数模型进行训练。我们还发布了GRAND-ROUNDS,这是一个新的大规模医生标注基准,包含来自七项研究中11位医生对9,217个评分的标注。我们表明,在典型的“LLM作为评审者”方法中,前沿LLM常常与医生意见相左,且彼此之间也存在分歧;但在少量案例上微调PrecepTron,能够使其在各项任务中实现与医生水平一致的评分。我们使用PrecepTron重现了发表在JAMA、Science和Nature Medicine上的五项评估LLM用于临床护理的有影响力研究的主要发现,而无需新的人工评分。随后,我们利用PrecepTron提出了关于LLM在医学中如何推理的新问题,这些问题仅靠人工评分是无法实现的,包括在前沿LLM被逐条甚至逐词提供临床病例时测量其诊断准确性。总之,PrecepTron和GRAND-ROUNDS为可重现的大规模研究LLM在医学中的推理方式奠定了基础。所有代码、数据和标签均免费提供给研究人员。

英文摘要

Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific questions investigated and makes it unclear whether findings would be reproduced with a different set of evaluators. To more rigorously and scalably study clinical reasoning in AI models, here we introduce PrecepTron, an LLM fine-tuned for physician-level evaluation of open-ended responses. PrecepTron was trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples. We also release GRAND-ROUNDS, a new large-scale physician-annotated benchmark of 9,217 scored responses from 160 clinicians across seven studies. We show that frontier LLMs in typical "LLM-as-a-judge" approaches often disagree with physicians and with each other, but fine-tuning PrecepTron on a small number of cases enables physician-level consistent scoring across tasks. We use PrecepTron to reproduce headline findings from five influential studies assessing LLMs for clinical care in JAMA, Science, and Nature Medicine without new human grading. Using PrecepTron, we then pose new questions about how LLMs reason in medicine that would have been infeasible with human grading alone, including measuring the diagnostic accuracy of frontier LLMs when clinical cases are provided piecemeal, even token by token. Together, PrecepTron and GRAND-ROUNDS provide a foundation for reproducible, large-scale study of how LLMs reason in medicine. All code, data, and labels are made freely available for researchers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑