arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估具有挑战性的真实世界临床病例中的多轮多模态诊断推理

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao, Jonathan Chong Kai Liew, Matthew Yu Heng Wong, Nicolás Lescano, Nikita R. Paripati, Emily Ling-Lin Pai, Jiarui Liu, Heli Qi, Heng-Jui Chang, Benny Kai Guo Loo, Huitao Li, Kunyu Yu, Yufan Wang, Chuan Hong, Shijian Lu, Douglas Teodoro, Naoto Yokoya, Ross Koppel, Mona Diab, Hua Xu, David W. Bates, Nan Liu, Yifan Peng

arXiv 2607.25933首次发表:更新:

发表机构

Center for Biomedical Data Science, Duke-NUS Medical School; Duke-NUS AI + Medical Sciences Initiative, Duke-NUS Medical School; Department of Population Health Sciences, Weill Cornell Medicine; System Engineering, College of Engineering, Cornell University; Graduate School of Frontier Sciences, The University of Tokyo; RIKEN Center for Advanced Intelligence Project; Department of Biostatistics and Bioinformatics, Duke University; Perelman School of Medicine, University of Pennsylvania; School of Clinical Medicine, University of Cambridge; Hospital of the University of Pennsylvania; Children’s Hospital of Philadelphia (CHOP); Department of Anatomic Pathology and Laboratory Medicine, Hospital of the University of Pennsylvania; Department of Pathology and Laboratory Medicine, University of California, San Francisco; Language Technologies Institute, Carnegie Mellon University(生物医学数据科学中心,杜克 - 新加坡国立大学医学院; 杜克 - 新加坡国立大学人工智能与医学科学计划,杜克 - 新加坡国立大学医学院; 人口健康科学系,威尔康奈尔医学院; 系统工程,康奈尔大学工程学院; 前沿科学研究生院,东京大学; 理化学研究所先进智能项目中心; 生物统计学与生物信息学系,杜克大学; 佩雷尔曼医学院,宾夕法尼亚大学; 临床医学学院,剑桥大学; 宾夕法尼亚大学医院; 费城儿童医院; 解剖病理学与检验医学系,宾夕法尼亚大学医院; 病理学与检验医学系,加州大学旧金山分校; 语言技术研究所,卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有多模态大语言模型评估难以捕捉现实临床诊断复杂性的问题,开发ClinMM-Bench基准,用两级框架评估15个模型,发现专有模型诊断准确性最高,但各模型完全正确诊断比例有限,当前模型诊断推理有局限并明确了失败模式。

AI 中文摘要

临床诊断评估不仅应评估模型能否提供正确诊断,还应反映临床实践的现实情况,包括多模态信息的逐步披露、诊断假设的动态更新以及临床推理的持续完善。然而,现有的多模态大语言模型评估通常依赖单轮或孤立任务,难以充分捕捉现实世界临床诊断的复杂性。为弥补这一差距,我们开发了ClinMM-Bench,这是迄今为止最大的多轮多模态临床诊断评估基准。ClinMM-Bench包含1089个具有挑战性的真实世界临床病例和八个专业的3760幅医学图像。我们使用两级评估框架系统地评估了15个代表性的多模态大语言模型,该框架同时评估诊断准确性和诊断推理质量。结果表明,专有模型总体诊断准确性最高,但所有模型中完全正确诊断的比例仍然有限。在诊断推理质量方面,当前模型可以识别合理的诊断方向,但在生成可靠的诊断推理方面仍有相当大的局限性。错误分析进一步确定了五种代表性的失败模式:信息合成失败、知识映射错误、感知错误、过早结束和视觉幻觉。

英文摘要

Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single-turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinical diagnosis. To bridge this gap, we developed ClinMM-Bench, the largest multi-turn multimodal clinical diagnostic evaluation benchmark to date. ClinMM-Bench contains 1,089 challenging real-world clinical cases and 3,760 medical images across eight specialties. We systematically evaluated 15 representative MLLMs using a two-level evaluation framework that assessed both diagnostic accuracy and diagnostic reasoning quality. Results showed that proprietary models achieved the highest overall diagnostic accuracy, but the proportion of completely correct diagnoses remained limited across all models. In terms of diagnostic reasoning quality, current models can identify plausible diagnostic directions but still have considerable limitations in generating reliable diagnostic reasoning. Error analysis further identified five representative failure modes: information synthesis failure, knowledge mapping error, perception error, premature closure, and visual hallucination.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑