MedGUIDE:基准测试大型语言模型的临床决策能力
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
- Harvard University(哈佛大学)
- MIT(麻省理工学院)
- Cornell University(康奈尔大学)
- Mayo Clinic(梅奥诊所)
- UIUC(伊利诺伊大学厄巴纳-香槟分校)
- University of Virginia(弗吉尼亚大学)
- Harvard Medical School(哈佛医学院)
- Abaka AI
- NYU(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出 MedGUIDE 基准,基于 55 棵 NCCN 癌症决策树和 7,747 个高质量样本评估 25 个 LLM 的指南遵循能力,揭示即使医学专用模型也存在明显不足。
AI中文摘要:
临床指南通常以决策树形式构建,是循证医学实践的核心,对于确保安全、准确的诊断决策至关重要。然而,目前尚不清楚大型语言模型(LLM)能否可靠遵循此类结构化规程。在本研究中,我们提出 MedGUIDE,这是一个用于评估 LLM 作出与指南一致的临床决策能力的新基准。MedGUIDE 基于覆盖 17 种癌症类型的 55 棵精选 NCCN 决策树构建,并利用 LLM 生成的临床场景创建大量多项选择题诊断问题。我们采用两阶段质量筛选流程,结合专家标注的奖励模型以及跨十项临床和语言标准的 LLM-as-a-judge 集成,最终选出 7,747 个高质量样本。我们评估了 25 个 LLM,涵盖通用模型、开源模型和医学专用模型,发现即使是特定领域 LLM,在需要遵循结构化指南的任务上也往往表现不佳。我们还测试了通过在上下文中纳入指南或持续预训练能否提升性能。研究结果强调了 MedGUIDE 的重要性,可用于评估 LLM 是否能在真实临床环境所要求的程序框架内安全运行。
英文摘要:
Clinical guidelines, typically structured as decision trees, are central to evidence-based medical practice and critical for ensuring safe and accurate diagnostic decision-making. However, it remains unclear whether Large Language Models (LLMs) can reliably follow such structured protocols. In this work, we introduce MedGUIDE, a new benchmark for evaluating LLMs on their ability to make guideline-consistent clinical decisions. MedGUIDE is constructed from 55 curated NCCN decision trees across 17 cancer types and uses clinical scenarios generated by LLMs to create a large pool of multiple-choice diagnostic questions. We apply a two-stage quality selection process, combining expert-labeled reward models and LLM-as-a-judge ensembles across ten clinical and linguistic criteria, to select 7,747 high-quality samples. We evaluate 25 LLMs spanning general-purpose, open-source, and medically specialized models, and find that even domain-specific LLMs often underperform on tasks requiring structured guideline adherence. We also test whether performance can be improved via in-context guideline inclusion or continued pretraining. Our findings underscore the importance of MedGUIDE in assessing whether LLMs can operate safely within the procedural frameworks expected in real-world clinical settings.