arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2405.20421cs.AI

比随机还差?一项对大型多模态模型在医学VQA中极其简单的探测评估

Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA

  • University of California, Santa Cruz(加州大学圣克ruz分校)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Qianqi Yan, Xuehai He, Xiang Yue, Xin Eric Wang

更新

AI总结:

提出 ProbMed 数据集,通过否定配对探测和多维度程序化诊断揭示主流 LMM 在医学影像细粒度诊断上甚至差于随机猜测。

AI中文摘要:

大型多模态模型(LMM)在医学视觉问答(Med-VQA)中已取得显著进展,并在现有基准上实现高准确率。然而,其在稳健评估下的可靠性仍存疑。本研究揭示,当接受简单的探测评估时,最先进模型在医学诊断问题上的表现比随机猜测还差。为解决这一关键评估问题,我们提出 Probing Evaluation for Medical Diagnosis(ProbMed)数据集,通过探测评估和程序化诊断严格评估 LMM 在医学影像中的表现。具体而言,探测评估的特点是将原始问题与带有幻觉属性的否定问题配对;程序化诊断则要求针对每张图像跨多个诊断维度进行推理,包括模态识别、器官识别、临床发现、异常和位置定位。我们的评估表明,GPT-4o、GPT-4V 和 Gemini Pro 等表现领先的模型在专门诊断问题上的表现比随机猜测还差,表明其在处理细粒度医学询问方面存在显著局限。此外,LLaVA-Med 等模型甚至在更一般的问题上也表现吃力,而 CheXagent 的结果证明了专业知识可在同一器官的不同模态间迁移,表明专门领域知识对于提升性能仍然至关重要。本研究强调,迫切需要更稳健的评估,以确保 LMM 在医学诊断等关键领域中的可靠性,而当前 LMM 仍远不能应用于这些领域。

英文摘要:

Large Multimodal Models (LMMs) have shown remarkable progress in medical Visual Question Answering (Med-VQA), achieving high accuracy on existing benchmarks. However, their reliability under robust evaluation is questionable. This study reveals that when subjected to simple probing evaluation, state-of-the-art models perform worse than random guessing on medical diagnosis questions. To address this critical evaluation problem, we introduce the Probing Evaluation for Medical Diagnosis (ProbMed) dataset to rigorously assess LMM performance in medical imaging through probing evaluation and procedural diagnosis. Particularly, probing evaluation features pairing original questions with negation questions with hallucinated attributes, while procedural diagnosis requires reasoning across various diagnostic dimensions for each image, including modality recognition, organ identification, clinical findings, abnormalities, and positional grounding. Our evaluation reveals that top-performing models like GPT-4o, GPT-4V, and Gemini Pro perform worse than random guessing on specialized diagnostic questions, indicating significant limitations in handling fine-grained medical inquiries. Besides, models like LLaVA-Med struggle even with more general questions, and results from CheXagent demonstrate the transferability of expertise across different modalities of the same organ, showing that specialized domain knowledge is still crucial for improving performance. This study underscores the urgent need for more robust evaluation to ensure the reliability of LMMs in critical fields like medical diagnosis, and current LMMs are still far from applicable to those fields.

↑