arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22323cs.CV

MedReaMM:评估大型多模态模型在专家级临床诊断综合能力上的表现

MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

Lai Wei, Yuchao Chen, Zhenbiao Cao, Xiaojin Zhang, Zhongyu Wei, Bangting Wang, Wei Chen, Xiang Bai

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建了多模态临床诊断基准MedReaMM,评估23个大型多模态模型的诊断综合能力,发现多数模型准确率不足50%,揭示了该领域的能力差距及相关影响因素。

中文摘要 AI 辅助

大型语言模型(LLMs)在诊断决策中的应用已受到越来越多的关注。然而,现有的基准大多聚焦于文本推理或孤立的视觉问答(VQA)任务,缺乏对临床叙事与医学影像的整体整合,因此无法评估专家临床判断核心的多模态诊断综合能力。为弥合这一差距,我们推出MedReaMM——一个专门用于评估模型在全信息范式下,将由详细患者病史与多张医学影像组成的异质性临床证据综合为准确鉴别诊断能力的基准。该基准由顶级医学期刊的病例报告与精心整理的临床病例数据库构建而成,包含625个经专家验证的病例,每个病例平均有2.79张医学影像,共标注了1042个符合ICD-11编码的标准化诊断。这些病例主要代表罕见、非典型或多系统表现,需要超越常规模式识别的专家级证据整合。我们评估了23个大型多模态模型(LMMs),发现大多数模型的诊断准确率低于50%,凸显出多模态诊断综合能力存在巨大差距。进一步分析表明,医学知识熟练度、医学影像理解能力与证据整合能力均与诊断表现高度相关。

英文摘要

The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus on textual reasoning or isolated visual question-answering (VQA) tasks, lacking holistic integration of clinical narratives and medical imaging, and thus failing to assess the multimodal diagnostic synthesis capability central to expert clinical judgment. To bridge this gap, we introduce MedReaMM, a benchmark specifically designed to evaluate models' ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm. Constructed from case reports sourced from top-tier medical journals and curated clinical case databases, MedReaMM comprises 625 expert-validated cases with an average of 2.79 medical images per case and a total of 1,042 standardized diagnoses annotated with ICD-11 codes. These cases predominantly represent rare, atypical, or multi-system presentations that demand expert-level evidence integration beyond routine pattern recognition. We evaluate 23 Large Multimodal Models (LMMs) and find that most achieve diagnostic accuracy scores below 50%, underscoring a substantial gap in multimodal diagnostic synthesis capability. Further analysis reveals that medical knowledge proficiency, medical image understanding, and evidence integration are all highly correlated with diagnostic performance.

发表机构

  • The First Affiliated Hospital with Nanjing Medical University(南京医科大学第一附属医院)
  • Jiangsu Province Hospital(江苏省医院)
  • Huazhong University of Science Technology(华中科技大学)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑