Are Large Vision Language Models Truly Grounded in Medical Images? Evidence from Italian Clinical Visual Question Answering
大视觉语言模型真的在医学图像上具有基础性吗?来自意大利临床视觉问答的证据
机构 * SIIAM ; NSBProject ; Dept. of Life Sciences & Public Health, UCSC(生命科学与公共卫生系,UCSC) ; ASL RM 4 ; UCSC ; Univ. Paris Cité(巴黎Cité大学) ; Univ. of Pisa(比萨大学)
专题命中 视觉问答 :vision language model(title,abstract);visual question answering(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
AI总结 研究通过测试四种先进模型在意大利医学问题上的表现,揭示了大视觉语言模型在视觉基础上的差异,强调了临床部署前的严格评估需求。
Comments Accepted at the Workshop on Multimodal Representation Learning for Healthcare (MMRL4H), EurIPS 2025