X-PCR:眼科诊断中跨模态渐进临床推理的基准测试
X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis
- School of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
- Tsinghua University(清华大学)
- University of Nottingham(诺丁汉大学)
- School of AI, Shenzhen University(深圳大学人工智能学院)
- Wenzhou Medical University(温州医科大学)
- Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出X-PCR基准测试,通过完整眼科诊断流程评估多模态大语言模型的渐进推理和跨模态整合能力,包含26,415张图像和177,868个专家验证的VQA对,评估21个模型揭示其在渐进推理和跨模态整合方面的不足。
AI中文摘要:
尽管多模态大语言模型(MLLMs)在多模态诊断中的临床推理能力取得显著进展,但其在多模态诊断中的临床推理能力仍缺乏充分评估。当前的基准测试大多基于单模态数据,无法评估临床实践中必需的渐进推理和跨模态整合能力。本文引入了跨模态渐进临床推理(X-PCR)基准测试,这是首个通过完整眼科诊断流程对MLLMs进行全面评估的基准测试,包含两个推理任务:1)一个跨越图像质量评估到临床决策的六阶段渐进推理链,2)一个整合六种成像模态的跨模态推理任务。该基准测试包含26,415张图像和177,868个专家验证的VQA对,涵盖52种眼科疾病。对21个MLLMs的评估揭示了渐进推理和跨模态整合方面的关键差距。数据集和代码:https://github.com/CVI-SZU/X-PCR。
英文摘要:
Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely unexamined. Current benchmarks, mostly single-modality data, can't evaluate progressive reasoning and cross-modal integration essential for clinical practice. We introduce the Cross-Modality Progressive Clinical Reasoning (X-PCR) benchmark, the first comprehensive evaluation of MLLMs through a complete ophthalmology diagnostic workflow, with two reasoning tasks: 1) a six-stage progressive reasoning chain spanning image quality assessment to clinical decision-making, and 2) a cross-modality reasoning task integrating six imaging modalities. The benchmark comprises 26,415 images and 177,868 expert-verified VQA pairs curated from 51 public datasets, covering 52 ophthalmic diseases. Evaluation of 21 MLLMs reveals critical gaps in progressive reasoning and cross-modal integration. Dataset and code: https://github.com/CVI-SZU/X-PCR.