arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.25559cs.AIcs.LG

放射科终极考核(RadLE):前沿多模态AI与人类专家的基准对比及放射学视觉推理错误分类体系

Radiology's Last Exam (RadLE): Benchmarking Frontier Multimodal AI Against Human Experts and a Taxonomy of Visual Reasoning Errors in Radiology

  • Centre for Responsible Autonomous Systems in Healthcare (CRASH) Lab, Koita Centre for Digital Health(负责任的自主医疗系统中心(CRASH)实验室、Koita数字健康中心)
  • Ashoka University(阿什oka大学)
  • Independent Researcher(独立研究者)
  • Koita Centre for Digital Health(Koita数字健康中心)

机构由 AI 辅助整理,请以论文原文为准。

Suvrankar Datta, Divya Buchireddygari, Lakshmi Vennela Chowdary Kaza, Mrudula Bhalke, Kautik Singh, Ayush Pandey, Sonit Sai Vasipalli, Upasana Karnwal, Hakikat … 展开作者

Suvrankar Datta, Divya Buchireddygari, Lakshmi Vennela Chowdary Kaza, Mrudula Bhalke, Kautik Singh, Ayush Pandey, Sonit Sai Vasipalli, Upasana Karnwal, Hakikat Bir Singh Bhatti, Bhavya Ratan Maroo, Sanjana Hebbar, Rahul Joseph, Gurkawal Kaur, Devyani Singh, Akhil V, Dheeksha Devasya Shama Prasad, Nishtha Mahajan, Ayinaparthi Arisha, Rajesh Vanagundi, Reet Nandy, Kartik Vuthoo, Snigdhaa Rajvanshi, Nikhileswar Kondaveeti, Suyash Gunjal, Rishabh Jain, Rajat Jain, Anurag Agrawal

更新

AI总结:

该研究构建含50个专家级疑难病例的RadLE基准,测试5款前沿多模态AI的放射诊断表现,发现AI准确率远低于放射科医师,同时提出视觉推理错误分类体系,为模型优化和评估标准制定提供参考。

AI中文摘要:

大语言模型(LLM)、视觉语言模型(VLM)等通用多模态AI系统正通过广泛可用的面向消费者的聊天机器人,越来越多地被临床医生和患者用于医学影像解读。多数声称达到专家级性能的评估,都是基于包含常见病症的公开数据集开展的。针对前沿模型在疑难诊断病例上表现的严格评估仍然十分有限。我们构建了一个包含50个跨多种影像模态的专家级“即时诊断”病例的试点基准,用以评估前沿AI模型与 board-certified 放射科医师、放射科受训人员的表现对比。为模拟真实使用场景,我们通过五款主流前沿AI模型的原生网页界面对其推理模式进行了测试,分别为OpenAI o3、OpenAI GPT-5、Gemini 2.5 Pro、Grok-4和Claude Opus 4.1。诊断准确率由设盲的专家评分,可重复性通过三次独立运行进行评估。我们还额外测试了GPT-5在不同推理模式下的表现。研究评估了推理质量错误,并定义了一套视觉推理错误分类体系。Board-certified 放射科医师取得了最高的诊断准确率(83%),优于受训人员(45%)和所有AI模型(表现最佳的GPT-5准确率为30%)。可靠性方面,GPT-5和o3为高水平,Gemini 2.5 Pro和Grok-4为中等水平,Claude Opus 4.1则为低水平。这些发现表明,先进的前沿模型在疑难诊断病例上的表现远不及放射科医师。我们的基准凸显了当前通用AI在医学影像领域的局限性,并警示不应在无监督情况下将其用于临床。我们还对推理轨迹进行了定性分析,提出了一套实用的AI模型视觉推理错误分类体系,以更好地理解其失效模式,为评估标准制定提供参考,并指导开发更鲁棒的模型。

英文摘要:

Generalist multimodal AI systems such as large language models (LLMs) and vision language models (VLMs) are increasingly accessed by clinicians and patients alike for medical image interpretation through widely available consumer-facing chatbots. Most evaluations claiming expert level performance are on public datasets containing common pathologies. Rigorous evaluation of frontier models on difficult diagnostic cases remains limited. We developed a pilot benchmark of 50 expert-level "spot diagnosis" cases across multiple imaging modalities to evaluate the performance of frontier AI models against board-certified radiologists and radiology trainees. To mirror real-world usage, the reasoning modes of five popular frontier AI models were tested through their native web interfaces, viz. OpenAI o3, OpenAI GPT-5, Gemini 2.5 Pro, Grok-4, and Claude Opus 4.1. Accuracy was scored by blinded experts, and reproducibility was assessed across three independent runs. GPT-5 was additionally evaluated across various reasoning modes. Reasoning quality errors were assessed and a taxonomy of visual reasoning errors was defined. Board-certified radiologists achieved the highest diagnostic accuracy (83%), outperforming trainees (45%) and all AI models (best performance shown by GPT-5: 30%). Reliability was substantial for GPT-5 and o3, moderate for Gemini 2.5 Pro and Grok-4, and poor for Claude Opus 4.1. These findings demonstrate that advanced frontier models fall far short of radiologists in challenging diagnostic cases. Our benchmark highlights the present limitations of generalist AI in medical imaging and cautions against unsupervised clinical use. We also provide a qualitative analysis of reasoning traces and propose a practical taxonomy of visual reasoning errors by AI models for better understanding their failure modes, informing evaluation standards and guiding more robust model development.

补充信息

↑