大型语言模型的临床决策路径遵循情况基准测试
Benchmarking Clinical Decision Pathway Adherence in Large Language Models
浏览论文内容
中文总结 AI 辅助
该研究针对现有医疗LLM基准缺乏指南遵循评估的问题,构建了MEGA-CDP基准,经实验发现当前LLM提供可靠临床决策支持仍具挑战,凸显了CDP导向评估的必要性。
中文摘要 AI 辅助
遵循临床实践指南定义的临床决策路径(CDP)对于安全可靠的医疗决策至关重要。然而,现有的医疗大型语言模型(LLM)基准主要评估最终答案的准确性,对模型遵循指南的能力评估有限。为解决这一缺口,我们推出MEGA-CDP,这是一个用于评估医疗LLM能否以提供的指南为参考生成符合指南的CDP的基准。MEGA-CDP由2274份中英文临床实践指南通过指南到病例的流程构建而成,得到42353个带有明确参考CDP的临床病例。它支持单轮 vignette 和多轮交互设置,并引入了面向CDP的评估框架以衡量路径一致性。对16个代表性LLM的实验表明,当前模型提供可靠临床决策支持仍具挑战性,凸显了面向CDP的评估的必要性,以及MEGA-CDP在推进医疗LLM指南遵循方面的价值。
英文摘要
Following clinical decision pathways (CDPs) defined by clinical practice guidelines is essential for safe and reliable medical decision-making. However, existing medical large language model (LLM) benchmarks mainly evaluate final-answer accuracy, providing limited evaluation of models' ability to adhere to guidelines. To address this gap, we introduce MEGA-CDP, a benchmark for evaluating whether medical LLMs can generate guideline-adherent CDPs using provided guidelines as references. MEGA-CDP is constructed from 2,274 English and Chinese clinical practice guidelines through a guideline-to-case pipeline, yielding 42,353 clinical cases with explicit reference CDPs. It supports both single-turn vignette and multi-turn interactive settings, and introduces a CDP-oriented evaluation framework for measuring pathway consistency. Experiments on 16 representative LLMs show that reliable clinical decision support remains challenging for current models, demonstrating the need for CDP-oriented evaluation and the value of MEGA-CDP for advancing guideline adherence in medical LLMs.
发表机构
- Tongji University(同济大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Tongji University School of Medicine(同济大学医学院)
- Sun Yat-sen University(中山大学)
- Fuzhou University(福州大学)
- Shenzhen Loop Area Institute(深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。