arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

医学检查表:评估多模态模型对医学图像的理解

Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models

Bannapol Limanond, Masanori Suganuma, Takayuki Okatani

arXiv 2607.21998首次发表:更新:

AI 中文总结

介绍医学检查表这一基准测试,通过给模型一张图像及一正一误两个标题进行二元测试,可统一评估不同多模态模型,验证其对医学概念的理解,评估其处理分布外输入的能力,用其评估四个先进模型发现它们理解图像能力待提升。

AI 中文摘要

本文介绍了一种用于评估医学多模态模型的新基准测试——医学检查表。多模态模型在医学视觉语言任务中展现出潜力,但评估其性能具有挑战性。医学检查表对模型进行二元测试,给定一张图像和两个标题(一个正确一个错误),错误标题中有一个医学概念被错误替换。该测试可统一评估不同多模态模型,验证其对各种医学概念的理解能力,减少数据偏差并评估处理分布外输入的能力。用其评估四个先进模型发现,它们在特定任务中表现出色,但可能无法正确理解图像,数据集和代码将在接受后公开。

英文摘要

This paper introduces a new benchmark test, Medical-Checklist, for assessing medical multimodal models. The recent advancements in multimodal models have demonstrated significant potential in the field of medical vision-language tasks. However, it is becoming increasingly clear that evaluating these models' performance, whether they are applied to natural or medical images, is challenging. The critical question is whether the models can accurately understand an input image while associating it with relevant input text. To address this, Medical-Checklist imposes a binary test on the models: they are given an image and two captions, where one is correct and the other incorrect, and the model must select the correct one. The incorrect caption contains a single medical concept (word or phrase) that is inaccurately substituted from the correct caption. Although the task is simple, this simplicity enables the unified assessment of diverse multimodal models designed and learned on different principles. It also enables us to verify whether models correctly understand a wide range of medical concepts across various medical sub-domains. Medical-Checklist is designed to reduce potential biases in data and to enable evaluation of the models' ability to handle out-of-distribution inputs, which were difficult in existing datasets. When evaluating four state-of-the-art medical multimodal models with Medical-Checklist, it was revealed that despite their excellent performance in specific tasks such as Med-VQA, they may not correctly understand images, suggesting a long journey ahead for clinical application. The dataset and code will be made public upon acceptance.

CommentsAccepted for publication in IEEE Journal of Biomedical and Health Informatics

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑