MTDiag:面向临床有意义的大语言模型评估的多轮诊断数据集
MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation
查看机构详情
- Technical University of Munich(慕尼黑工业大学)
- UT Southwestern(西南大学(德克萨斯州))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究针对现有医学大语言模型评估基准无法反映临床多轮交互诊断能力的问题,构建了多轮诊断数据集MTDiag,并提出临床知识导向的多轮鉴别诊断评估指标。
中文摘要 AI 辅助
临床诊断本质上是交互且渐进的,但评估医学领域大语言模型(LLMs)的主流范式仍是静态问答基准或基于模板的对话。这些基准几乎无法说明模型能否在动态临床场景中作为诊断智能体发挥作用,且LLMs在多轮场景中会出现显著的准确性和可靠性下降。为解决该问题,我们提出MTDiag,这是一个由DDXPlus、MIMIC-IV和已发表病例报告(AJCR)三个异构来源构建的大型多轮诊断对话数据集,涵盖常见急诊(ED)病症以及长尾罕见和非典型病症。所有病例均被归一化为规范模式,该模式锚定在最全面且广泛使用的医学知识库(UMLS概念标识符、ICD-10诊断编码)中。我们发布了该模式、基于UserLM-8B的话语生成流水线,以及由医生验证的将结构化临床证据转换为自然语言话语的数据集。重要的是,我们引入并推动了以临床知识为基础的指标,用于评估作为诊断智能体的LLMs,超出了诊断准确性的范畴,适用于多轮鉴别诊断任务。
英文摘要
Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say little about whether a model can serve as a diagnostic agent in a dynamic clinical encounter, with LLMs showing significant accuracy and reliability degradation in multi-turn settings. To address this issue, we present MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports (AJCR), covering common ED presentations as well as long-tail rare and atypical conditions. All cases are normalized into a canonical schema anchored in the most comprehensive and widely-adopted medical knowledge bases (UMLS concept identifiers, with ICD-10 diagnosis codes). We release the schema, a UserLM-8B-based utterance-generation pipeline, and the physician-validated dataset that converts structured clinical evidence into natural-language utterances. Importantly, we introduce and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.