arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

波兰医学视觉问答:视觉语言模型未充分利用视觉证据

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa

arXiv 2608.12928首次发表:更新:

发表机构

ARAAI Poland; NASK National Research Institute; Poznań University of Medical Sciences; T. Marciniak Lower Silesian Specialist Hospital; Adam Mickiewicz University(波兰ARAAI; NASK国家研究院; 波兹南医科大学; T.马尔奇尼亚克下西里西亚专科医院; 亚当·密茨凯维奇大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建了波兰医学VQA基准,评估多类视觉语言模型,发现模型未充分利用视觉证据,仅从文本或答案选项即可达到高于随机水平的准确率,仅GPT-5.6在部分子集上超人类表现。

AI 中文摘要

我们推出了一个波兰语医学视觉问答(VQA)基准,该基准由波兰医师和牙医专科认证委员会考试题目构建而成。此基准包含涵盖不同医学专业和视觉领域的含图像题目,以及仅含文本的问答(QA)对照组。我们评估了面向波兰语的通用开源权重和商业视觉语言模型。该任务仍具挑战性:最佳模型在完整VQA集上的准确率为79.0%,且仅GPT-5.6在有可用候选答案的子集上超过近似人类参考水平;所有其他被评估模型的表现均差于人类。为评估视觉定位,我们将完整输入与省略图像、问题或两者的配置进行对比,并按图像重要性对问题分类。模型从问题文本中获取的有用信息多于从图像中获取的信息,且在图像主导的问题上表现更差。尽管如此,在QA和VQA任务中,它们仅从答案选项就能达到高于随机水平的准确率,表明即使缺少关键任务组件,仍能保持非平凡的性能。

英文摘要

We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set. We evaluate Polish-oriented, general-purpose open-weight, and commercial vision-language models. The task remains challenging: the best model achieves 79.0\% accuracy on the full VQA set, and only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses; all other evaluated models perform worse than humans. To assess visual grounding, we compare complete inputs with configurations omitting the image, the question, or both, and categorize questions by image importance. Models derive more useful information from the question text than from the image and perform worse on image-dominant questions. Across both QA and VQA, they nevertheless achieve above-chance accuracy from the answer choices alone, showing that non-trivial performance can persist even when key task components are missing.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑